Spec2Game: Can LLMs Generate Complete Playable Games from Detailed Specifications?

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models can translate natural language specifications into executable and behaviorally correct interactive programs. To this end, we propose Spec2Game, a benchmark requiring models to generate Pygame-based game projects, accompanied by fine-grained test cases spanning rule variations and complexity levels. For evaluation, we integrate multi-source evidence from source code, runtime execution, and visual output to enable four-dimensional automated assessment. Experimental results reveal that while current models achieve high code executability and perform well in element modeling, they exhibit significant deficiencies in faithfully implementing rule mechanisms and termination logic as specified. These findings highlight the critical insight that high executability does not equate to correct implementation.
📝 Abstract
Generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language. To evaluate large language models'ability to realize detailed specifications as complete interactive programs, we introduce Spec2Game, a benchmark that requires models to generate complete Pygame projects from detailed natural-language game specifications. Spec2Game comprises 15 game families and 150 task instances, with one canonical task and nine controlled rule variants per family, spanning three levels of implementation complexity. Using source-code, runtime, and visual evidence, we evaluate generated projects along four dimensions---Executability, Specification Realization, Code Quality, and User-Facing Quality. Across 14 LLMs and 3,330 generated projects, we find that high executability does not imply faithful specification realization. Component-level analysis further shows that models perform substantially better on Game Element Modeling than on Rule and Mechanism Modeling or Goal and Termination Modeling, indicating that faithfully implementing game rules and termination logic remains a major challenge.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Code Generation
Specification Realization
Game Development
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spec2Game benchmark
LLM code generation
Pygame
specification realization
multi-dimensional evaluation
🔎 Similar Papers
2024-07-24IEEE Transactions on GamesCitations: 0
Y
Yixue Cai
The Chinese University of Hong Kong
Y
Yuzhe Zhao
Nankai University
H
Hanxiang Chao
Wuhan University
Q
Qingsen Ma
The Chinese University of Hong Kong
Z
Ziheng Xiong
The Chinese University of Hong Kong
Jinhu Qi
Jinhu Qi
PhD candidate in CUHK CSE
Agentic AILLMsReasoning
Irwin King
Irwin King
The Chinese University of Hong Kong
social computingmachine learningAIgraph neural networksNLP