RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of iterative overfitting, poor generalization, and difficulty surpassing playable baselines when large language models generate games. To this end, we propose a recursive self-improvement framework for autonomous agents that introduces a novel dual-loop mechanism. The inner loop enables targeted refinement through local exploration-based diagnosis, while the outer loop tracks global quality via a dynamic test suite. Furthermore, reinforcement learning is employed to internalize successful experiences into model parameters, thereby enhancing generative capabilities. Experimental results on GameCraft-Bench demonstrate that our approach significantly improves game generation quality. Notably, Qwen-based models outperform GPT-5.5 under this framework while reducing token consumption by an order of magnitude (11×).
📝 Abstract
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Problem

Research questions and friction points this paper is trying to address.

automatic game generation
iterative refinement
overfitting
game quality improvement
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-improvement
Autonomous Agentic Framework
Experience Internalization
Dual-loop Mechanism
Automatic Game Generation
🔎 Similar Papers