🤖 AI Summary
This study addresses the limitation of existing agent evaluation benchmarks that focus solely on final outcomes while neglecting exploration processes within stateful environments. We propose a novel evaluation framework built upon Roblox Studio, which uniquely incorporates exploratory behavior as an assessment dimension. By decoupling observation and editing tools, designing an eight-tool action space, and implementing executable verification scripts, the framework supports large-scale parallel evaluation. Experimental assessments across thirteen frontier models yield a best single-turn pass rate of 51.7%, demonstrating a significant positive correlation between exploration depth and task success rates. This work fills a critical gap in procedural evaluation metrics for autonomous agents, and the complete task suite has been open-sourced to facilitate further research.
📝 Abstract
We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task.
The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks.
Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks.
We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at https://github.com/Roblox/open-game-eval.