A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks in evaluating the ability of coding agents to faithfully implement interdependent requirements spanning logic, rendering, and interaction within long-form game design documents (GDDs). To this end, we construct a benchmark comprising 100 long GDDs and propose a dependency-aware contract-based evaluation framework. This framework translates requirements into formal contracts and integrates static code analysis with adaptive testing, enabling consistent cross-agent comparisons under fixed contractual specifications. Our experiments reveal that current agents struggle to jointly satisfy interdependent requirements. Furthermore, introducing a requirement-specific feedback mechanism yields a 10.9% relative improvement in GDD fidelity compared to a self-revision baseline.
📝 Abstract
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
Problem

Research questions and friction points this paper is trying to address.

Coding Agents
Game Design Documents
Benchmark
End-to-End Game Development
Requirement Faithfulness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Game Design Documents
Coding Agents
Dependency-aware Contract
Adaptive Playtesting
Benchmark
🔎 Similar Papers
2024-07-24IEEE Transactions on GamesCitations: 0