星际争霸:母巢之战基准测试
Brood War Bench

原始链接: https://bw.swerdlow.dev/report

本·斯沃德洛(Ben Swerdlow)的“母巢之战基准”(Brood War Bench)实验评估了各类人工智能模型在即时战略游戏《星际争霸:母巢之战》中的表现。 **主要研究结果如下:** * **表现水平:** 没有模型能达到初学者以上的水平。即使是最优秀的智能体,在构建复杂军队、防御简单攻击或执行连贯策略等基本任务上也显得力不从心。 * **表现最佳者:** Codex Astra 脱颖而出,表现明显优于其他模型。其成功往往归功于破坏性的“非正统战术”(cheese tactics)——例如派遣单个探机去骚扰敌方建筑,而非依靠精妙的宏观管理。 * **模型弱点:** 许多模型难以维持持续生产,往往会产生不协调的子智能体,导致无法协同作战。旧模型在游戏时常表现得如同在玩回合制游戏,在“思考”时处于停滞状态;而新模型则表现出对游戏时间成本的更好认知。尤其是 Grok,经常无法生产军队,将过多资源用于逻辑推理而非实际操作。 * **参与度:** Claude Fable 展现出了令人期待的雄心,经常尝试攀升科技树并“正统”地进行游戏,尽管其执行力不足以持续获胜。 总而言之,虽然这些模型目前还不是合格的玩家,但该基准测试突显了人工智能在即时战略环境中巨大的未来发展潜力。

Hacker News 上的一场讨论重点介绍了“Brood War Bench”项目。该项目展示了一款能够游玩经典策略游戏《星际争霸:母巢之战》的 AI 智能体。 社区成员对该基准测试表示赞赏,因为它关注长期策略、战术编排和平衡性——这些要素在简单的 AI 测试环境中常被忽视。讨论还探讨了 AI 训练的潜力,用户们探究智能体是否能通过分析历史录像文件或社区讨论来学习。尽管出现了关于《星际争霸》录像解析器可用性的技术询问,但该项目仍作为一项复杂的测试脱颖而出,验证了智能体在高级即时战略游戏中进行复杂决策的能力。
相关文章

原文

Which model wins at Brood War?

Key takeaways

  • None of the models played beyond a beginner level.
  • Codex Astra is the clear leader beating all other models consistently.
  • Grok models are not smart enough to play Brood War yet.
  • Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking.

Brood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn't done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.

01

Codex found cheese before it found macro

Codex's strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.

The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.

I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn't communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build.

This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing. In games where I helped direct Codex, it was much better at planning those moments and getting its subagents to work together.

The persistence was real. In G009, after losing its army and main base, Codex 5.6 Terra / medium lifted its last Command Center and moved it toward the opposite corner. It survived for another six minutes.

Six Probes cross the map
A Probe first, then Zealots in drips
The last Command Center runs

02

Grok spent the game between actions

Grok 4.6 frequently produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit.

The actions it did take rarely developed into a working control loop. In G003, Grok / xhigh made three Marines and never reached the enemy base. In G002, Grok / medium made two Zealots and also never crossed the map. These looked less like bad strategies than failures to keep observing and acting.

Forty-three minutes, no army

03

Fable earnestly tried to play the game

I found myself rooting for Claude Fable in more than a few games. Fable usually tried to build an economy and climb the tech tree instead of stopping at the first unit available. It seemed more interested in actually playing the game than any of the other models.

In G007 it reached a Lair, Spire, and Mutalisks and won. In G027 it added a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives before winning. Ambition did not guarantee execution: in G036 Fable reached a Factory and Academy but Opus 5 overran it.

Fable gets Mutalisks
Fable keeps climbing
The build does not become an army

No agent here played beyond beginner level

Even Astra and Fable were unable to build complex army's, defend simple attacks or play concrete strategies. A beginner playing photon rush would win every single one of these games.

That said, watching the agents play made me more excited than I have been in a while. This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do. I look forward to watching them get there.

Choose up to five

5:00 / 15:00

Technology investment

Completed research + upgrade levels

00.511.520:005:0010:0015:00

Codex Astra0.3Claude Fable0.2Grok 4.60

Workers

Completed workers alive

081624320:005:0010:0015:00

Codex Astra14.7Claude Fable16.7Grok 4.67.4

Army size

Completed army and support units

0204060800:005:0010:0015:00

Codex Astra8.3Claude Fable7.6Grok 4.62

Structures

Completed buildings, including add-ons

071421280:005:0010:0015:00

Codex Astra6.2Claude Fable7.4Grok 4.64.5

Minerals in the bank

Unspent minerals, not income

05001,0001,5002,0000:005:0010:0015:00

Codex Astra240.6Claude Fable248.6Grok 4.6536.6

Gas in the bank

Unspent gas, not income

05001,0001,5002,0000:005:0010:0015:00

Codex Astra151Claude Fable326.9Grok 4.6244.5

Supply used

Includes production in progress

03060901200:005:0010:0015:00

Codex Astra26.5Claude Fable26.6Grok 4.611.3

When games ended

Share of games ending per 5-minute window

Codex Astra: median 8:10. 0:00 to before 5:00: 3.9%; 5:00 to before 10:00: 54.9%; 10:00 to before 15:00: 31.4%; 15:00 to before 20:00: 5.9%; 25:00 to before 30:00: 2%; 30:00 to before 35:00: 2%. Claude Fable: median 10:37. 5:00 to before 10:00: 38.9%; 10:00 to before 15:00: 38.9%; 15:00 to before 20:00: 22.2%. Grok 4.6: median 9:15. 0:00 to before 5:00: 2%; 5:00 to before 10:00: 51%; 10:00 to before 15:00: 23.5%; 15:00 to before 20:00: 13.7%; 20:00 to before 25:00: 2%; 40:00 to before 45:00: 7.8%.Codex AstraMedian 8:1060%Codex Astra: 3.9% (2 games) ended from 0:00 to before 5:00Codex Astra: 54.9% (28 games) ended from 5:00 to before 10:00Codex Astra: 31.4% (16 games) ended from 10:00 to before 15:00Codex Astra: 5.9% (3 games) ended from 15:00 to before 20:00Codex Astra: 0% (0 games) ended from 20:00 to before 25:00Codex Astra: 2% (1 games) ended from 25:00 to before 30:00Codex Astra: 2% (1 games) ended from 30:00 to before 35:00Codex Astra: 0% (0 games) ended from 35:00 to before 40:00Codex Astra: 0% (0 games) ended from 40:00 to before 45:00Claude FableMedian 10:3760%Claude Fable: 0% (0 games) ended from 0:00 to before 5:00Claude Fable: 38.9% (7 games) ended from 5:00 to before 10:00Claude Fable: 38.9% (7 games) ended from 10:00 to before 15:00Claude Fable: 22.2% (4 games) ended from 15:00 to before 20:00Claude Fable: 0% (0 games) ended from 20:00 to before 25:00Claude Fable: 0% (0 games) ended from 25:00 to before 30:00Claude Fable: 0% (0 games) ended from 30:00 to before 35:00Claude Fable: 0% (0 games) ended from 35:00 to before 40:00Claude Fable: 0% (0 games) ended from 40:00 to before 45:00Grok 4.6Median 9:1560%Grok 4.6: 2% (1 games) ended from 0:00 to before 5:00Grok 4.6: 51% (26 games) ended from 5:00 to before 10:00Grok 4.6: 23.5% (12 games) ended from 10:00 to before 15:00Grok 4.6: 13.7% (7 games) ended from 15:00 to before 20:00Grok 4.6: 2% (1 games) ended from 20:00 to before 25:00Grok 4.6: 0% (0 games) ended from 25:00 to before 30:00Grok 4.6: 0% (0 games) ended from 30:00 to before 35:00Grok 4.6: 0% (0 games) ended from 35:00 to before 40:00Grok 4.6: 7.8% (4 games) ended from 40:00 to before 45:000:0015:0030:0045:00

Time-series charts show means of recorded player-runs at each game time. Finished games drop out; missing samples are not filled. Models pool their effort settings. Units and buildings count only once completed; army excludes workers, Overlords, eggs, larvae, and ammunition.

Win rate vs. cost

Average cost per game, using the same prices as the leaderboard. Codex and Sonnet costs are token-based estimates.

Win rate0%25%50%75%100%$0.1$0.5$1$5$10$20Cost per game (USD, log scale)

How the benchmark ran

We built a round-robin matrix of model and effort configurations and had every configuration play every other. The harness ran those matchups in parallel across Freestyle VMs, saving game-engine data and both agents' harness logs for each match.

Head-to-head matrix

Read across a row. W is a win, L is a loss, and T is a match that reached the benchmark time limit.

Open the full 19 × 19 matrix

W win L loss T time limit

System12345678910111213141516171819
1Codex Astra / xhigh-WWWWWWWWWWWWWWWWWW
2Codex Astra / mediumL-WWWWWWWWWWLWWWWWW
3Codex Astra / lowLL-WWWWWWWWWLLWWWWW
4Codex 5.6 Sol / xhighLLL-LWWWLWWWLLWWWWW
5Codex 5.6 Sol / mediumLLLW-WWWLWWWLWWWWWW
6Codex 5.6 Sol / lowLLLLL-WWWWWWWWLWWWW
7Codex 5.6 Luna / xhighLLLLLL-WLWLWLLLWWWW
8Codex 5.6 Luna / mediumLLLLLLL-LLLLLWWWWWW
9Codex 5.6 Luna / lowLLLWWLWW-WLLLLLWWWW
10Codex 5.6 Terra / xhighLLLLLLLWL-WWLWWWWWW
11Codex 5.6 Terra / mediumLLLLLLWWWL-LLLWWWWW
12Codex 5.6 Terra / lowLLLLLLLWWLW-LLWWWWW
13Claude FableLWWWWLWWWWWW-LWWWWW
14Claude Opus 5LLWWLLWLWLWWW-WWWWW
15Claude SonnetLLLLLWWLWLLLLL-WWWW
16Claude HaikuLLLLLLLLLLLLLLL-LTT
17Grok 4.6 / xhighLLLLLLLLLLLLLLLW-WT
18Grok 4.6 / mediumLLLLLLLLLLLLLLLTL-W
19Grok 4.6 / lowLLLLLLLLLLLLLLLTTL-

Play your own match

Bring your agent and play Brood War with friends.

Play Brood War
联系我们 contact @ memedata.com