GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Brood War Bench: Codex Astra Goes 18-0, Grok Spent 43 Minutes Thinking and Never Built an Army

Ben Swerdlow ran 171 matches of StarCraft: Brood War using only AI agents as players and published the results at bw.swerdlow.dev. The benchmark was not designed as a rigorous capability test — Swerdlow built the playable agent interface to play with friends and then wondered how far models could go unsupported. The answer is: nowhere near the frontier of competitive play, but the failure modes are different enough per model family to be instructive.

Leaderboard

RankSystemW-LWin RateAPMCost/game
1Codex Astra / xhigh18-0100%12.6$10.54
2Codex Astra / medium16-288.9%17.2$15.11
3Claude Fable / standard15-383.3%12.6$12.24
4Codex Astra / low14-477.8%25.7$21.07
5Codex 5.6 Sol / medium13-572.2%10.1$5.12
6Codex 5.6 Sol / low12-666.7%18.1$9.23
7Claude Opus 512-666.7%10.5$20.78
8Codex 5.6 Sol / xhigh11-761.1%8.0$3.23
9Codex 5.6 Luna / low9-950.0%23.8$0.42
10Codex 5.6 Terra / xhigh9-950.0%15.8$2.10

Grok 4.6 is not on this table. That requires explanation.

What Happened to Grok

In G043, Grok 4.6 / xhigh logged 11,138 reasoning tokens across a 43-minute game. It issued six command batches. It never fielded a combat unit.

This was not an anomaly. G003: Grok / xhigh built three Marines and never crossed the map. G002: Grok / medium produced two Zealots, also never crossed the map. The pattern Swerdlow describes is not a bad strategy — it is a failure to maintain the observe-act loop at the game’s tempo. Brood War runs in real time. Agents that spend 40 seconds deliberating between commands are losing workers to the enemy while thinking.

The irony is that xhigh effort performs worse than lower-effort settings on this task. Grok’s failure mode is reasoning volume exceeding action cadence — a problem that compounds when the model is given more thinking budget.

Codex Found Cheese Before It Found Macro

Codex Astra’s winning approach was not superior strategy. It was disruption. Playing Protoss, Codex frequently sent a probe across the map to attack enemy workers or buildings before any army existed. The probe harass worked because opposing agents often spent dozens of seconds deciding how to respond to a single probe — time during which their economy stalled.

Codex was substantially weaker at the part of the game that follows a probe harass: sustained production, tech advancement, army timing. The model trickled units into defended bases and sent them in one at a time rather than waiting for critical mass. Swerdlow notes that Codex frequently created separate subagents for economy, army production, and army control — but the subagents communicated poorly. The army subagent would send each new unit forward immediately, unaware that the production subagent was building toward a larger planned attack.

This is a recognizable multi-agent coordination failure. Isolated subagents optimizing local objectives without shared state.

Fable Tried to Play the Actual Game

Claude Fable’s playstyle stands out qualitatively. It was the only model that consistently attempted to build a real economy and climb the technology tree rather than stopping at the first available combat unit.

In G007 it reached a Lair, Spire, and Mutalisk — mid-tier Zerg air units. In G027 it built a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives — the full Protoss tech stack — before winning. Ambition did not guarantee execution: in G036 it had a Factory and Academy in production when Claude Opus 5 overran the base.

Fable’s APM matches Codex Astra at 12.6, but the composition of those actions is different. Fable spends more commands on economic development.

What It Means

No agent here played beyond a beginner level, including the undefeated one. Codex Astra wins because its probe harass is cheap to execute and devastatingly effective against agents that freeze on edge cases. In a match against a competent human player, the same strategy would be scouted and answered in seconds.

The more useful signal from this benchmark is the failure taxonomy. Models that over-reason on a real-time task lose not because their reasoning is wrong but because the game does not pause while they think. That is not a Brood War problem — it is an exact model for any agentic deployment with a real-time feedback loop: infrastructure incident response, live trading, interactive robotics. Grok’s 11,138-token deliberation in a 43-minute game with six actions is a concrete illustration of what happens when a model’s thinking-to-acting ratio is calibrated for a different problem class.

The cost column also warrants attention. Codex 5.6 Luna / low at $0.42 per game goes 9-9 — 50% win rate. Codex Astra / xhigh at $10.54 per game goes 18-0. The performance gap is real. So is the 25x cost difference. For applications where throughput matters, the cost-win-rate curve matters more than the top of the leaderboard.