SWE-Game: Can Coding Agents Build the Games We Want?

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inadequacy of existing benchmarks in evaluating coding agents’ ability to construct playable games. We propose SWE-Game, the first end-to-end development benchmark grounded in real executable games, comprising 41 Godot reference games and 247 tasks spanning the entire pipeline from brief generation to engine porting. The benchmark enables independent verification through shared instrumented interfaces and establishes a multidimensional evaluation framework integrating runtime checks, behavioral replay, and VLM-based visual scoring. Experimental results demonstrate that Opus5 achieves the best performance yet attains a construction score below 60. Runtime checks reach 92.59% accuracy, significantly outperforming video-based VLM judging, while visual scoring exhibits a correlation coefficient of 0.829 with human evaluations.
πŸ“ Abstract
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.
Problem

Research questions and friction points this paper is trying to address.

coding agents
game development
benchmark evaluation
automated testing
playability assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Game Development Benchmark
Coding Agents Evaluation
Shared Instrumentation Interface
Vision-Language Rubrics
Executable State Checks
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Xiaoyu Chen
Xiaoyu Chen
Shanghai University
Cultural heritage informaticsHuman information behaviourSocial informaticsSocial impacts of AI
L
Lai Wei
Shanghai Jiao Tong University, Zhongguancun Academy
J
Jin Wang
Shanghai Jiao Tong University
X
Xiangyu Zou
Shenzhen University
R
Ruochen Fan
Shanghai Jiao Tong University
E
Enze Luo
Shanghai Jiao Tong University
M
Mingzhe Yao
Shanghai Jiao Tong University
J
Jiahui Zhu
Elbetech Technology
Y
Yuhua Wen
Beijing University of Posts and Telecommunications
Linghe Kong
Linghe Kong
Shanghai Jiao Tong University
Internet of ThingsMobile computingBig data
W
Weiran Huang
Shanghai Jiao Tong University, Shanghai Innovation Institute