Coding agents building Godot games from briefs score at most 50/100
Claude Opus 5 leads all five task types
Confirmed
The takeaway
SWE-Game (arXiv 2609.33678, Sept 27, 2026) tests coding agents on 247 game-development tasks grounded in 41 executable Godot reference games.
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 2 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 Astra150 days after its cutoff
- Claude Opus 5.589 days after its cutoff
- Gemini 3.8 Flash180 days after its cutoff
- Grok 4.7119 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 89 days before it.
Key facts
- 247 tasks over 41 executable Godot reference games in 13 gameplay categories (2D and 3D)
- Five task types: brief-to-game, implementation from a game design document, skeleton completion, repair of 83 injected faults, Godot-to-Unity porting
- Six models evaluated; Claude Opus 5 (‘Opus5’) best; best construction scores below 60/100; Brief-to-Game best 50.38
- Main failure modes: requirement omissions and gameplay-logic errors
- Evaluator validation: executable checks 92.59% balanced accuracy on human-labelled behaviours from 100 agent-built games vs 78.41% for a video-based VLM judge; visual rubric scores 0.829 Spearman with human raters
- Authors: Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang; v3 posted Oct 7, 2026
What happened
A research team released SWE-Game, a benchmark that asks coding agents to build, finish, repair and port real games in the Godot engine. It checks results by running the games and asserting on their behaviour, and by scoring screenshots against a rubric. Claude Opus 5 led, but the best agent still scored only about half marks when asked to build a game from a short brief.
Why it matters
Coding agents are now judged mostly on repository bug fixes (SWE-bench). Games test whether an agent can turn loose requirements into a working interactive program, and the low scores show a large gap. This matters for the wave of AI-built game clones seen in autumn 2026.
Sources
2 sources from 2 sites. Numbers match the chips in the text.
2 sources: 1 primary, 1 press
Primary
Press
- AI Weekly: SWE-Game, Opus 5 leads 247 Godot agent tasks, still below 60aiweekly.co, press
Changes
- Filed (explainx.ai Oct 9 digest); paper read on arXiv