{"schema":"postcutoff/event@1","as_of":"2026-10-10T23:43:00+02:00","url":"https://postcutoff.com/e/2026-09-27-swe-game-benchmark-godot/","md":"https://postcutoff.com/e/2026-09-27-swe-game-benchmark-godot/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"2026-09-27-swe-game-benchmark-godot","date":"2026-09-27","date_precision":"day","short_title":"SWE-Game benchmark","deck":"Coding agents building Godot games from briefs score at most 50/100; Claude Opus 5 leads all five task types","takeaway":"SWE-Game (arXiv 2609.33678, Sept 27, 2026) tests coding agents on 247 game-development tasks grounded in 41 executable Godot reference games.","category":"benchmark","category_label":"Benchmarks","importance":2,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":[{"n":1,"title":"arXiv 2609.33678: SWE-Game, Can Coding Agents Build the Games We Want?","url":"https://arxiv.org/abs/2609.33678","type":"paper","group":"primary","domain":"arxiv.org"},{"n":2,"title":"AI Weekly: SWE-Game, Opus 5 leads 247 Godot agent tasks, still below 60","url":"https://aiweekly.co/alerts/swe-game-opus5-leads-247-godot-agent-tasks-still-below-60","type":"press","group":"press","domain":"aiweekly.co"}],"official":1,"filed":"2026-10-10","updated":"2026-10-10","orgs":["SWE-Game authors"],"title":"SWE-Game benchmark: coding agents building Godot games from briefs score at most 50/100; Claude Opus 5 leads all five task types","summary":"SWE-Game (arXiv 2609.33678, Sept 27, 2026) tests coding agents on 247 game-development tasks grounded in 41 executable Godot reference games. Of six models, Claude Opus 5 was best on every task type, but no model scored above 60/100 on building games, and turning a short brief into a game topped out at 50.38. Executable runtime checks judged agent-built games far more accurately than a video-watching VLM judge.","key_facts":["247 tasks over 41 executable Godot reference games in 13 gameplay categories (2D and 3D)","Five task types: brief-to-game, implementation from a game design document, skeleton completion, repair of 83 injected faults, Godot-to-Unity porting","Six models evaluated; Claude Opus 5 ('Opus5') best; best construction scores below 60/100; Brief-to-Game best 50.38","Main failure modes: requirement omissions and gameplay-logic errors","Evaluator validation: executable checks 92.59% balanced accuracy on human-labelled behaviours from 100 agent-built games vs 78.41% for a video-based VLM judge; visual rubric scores 0.829 Spearman with human raters","Authors: Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang; v3 posted Oct 7, 2026"],"key_numbers":[],"tags":["benchmark","coding-agents","games","godot","unity","evaluation"],"science":null,"body_md":"## What happened\n\nA research team released SWE-Game, a benchmark that asks coding agents to build, finish, repair and port real games in the Godot engine.\nIt checks results by running the games and asserting on their behaviour, and by scoring screenshots against a rubric.\nClaude Opus 5 led, but the best agent still scored only about half marks when asked to build a game from a short brief.\n\n## Why it matters\n\nCoding agents are now judged mostly on repository bug fixes (SWE-bench). Games test whether an agent can turn loose requirements into a\nworking interactive program, and the low scores show a large gap. This matters for the wave of AI-built game clones seen in autumn 2026.","disputed":[],"related":[{"id":"2026-07-24-claude-opus-5","url":"https://postcutoff.com/e/2026-07-24-claude-opus-5/","date":"2026-07-24","date_precision":"day","short_title":"Anthropic releases Claude Opus 5","deck":"Near-Fable-5 intelligence at half the price","takeaway":"Opus 5 brought most of Fable 5's capability to half the price.","category":"model-release","category_label":"Model releases","importance":4,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":8,"official":3,"filed":"2026-09-29","updated":"2026-09-29","orgs":["Anthropic"]}],"people":[],"posts":[],"videos":[],"models":[],"changes":[{"date":"2026-10-10","type":"filed","text":"Created (explainx.ai Oct 9 digest); paper read on arXiv"}],"provenance":{"agents":[{"model":"Claude Opus 5.5","maker":"Anthropic","tool":"Claude Code"}],"filed":"2026-10-10","run":null,"sources_read":"(explainx.ai Oct 9 digest); paper read on arXiv","updated":"2026-10-10","human_review":null,"version":null},"gaps":[{"model_id":"gpt-6-astra","name":"GPT-6 Astra","cutoff":"2026-04","days_after":150,"in_training_data":false},{"model_id":"claude-opus-5-5","name":"Claude Opus 5.5","cutoff":"2026-06","days_after":89,"in_training_data":false},{"model_id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","cutoff":"2026-03","days_after":180,"in_training_data":false},{"model_id":"grok-4-7","name":"Grok 4.7","cutoff":"2026-05","days_after":119,"in_training_data":false}],"short_url":null}