🤖 AI Summary
This study addresses the limitation that AI-generated game code, while functionally correct, often lacks playability. To bridge this gap, we propose an experience-oriented recursive development framework featuring a closed-loop multi-agent architecture comprising designer, builder, player, and reviewer agents. Notably, this work introduces a code-native player mechanism that efficiently collects trajectories via programmatic interfaces to eliminate GUI-based evaluation biases, alongside a trajectory-based preference induction metric for iterative optimization. Experimental results demonstrate that the proposed method achieves a score of 77.89 on GameCraft-Bench and improves the success rate by 34.1% on GameASG-Bench. Furthermore, user playtime and ratings increase significantly, effectively advancing AI-generated games from functional prototypes to genuinely entertaining experiences.
📝 Abstract
Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.