OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing agent evaluation benchmarks that focus solely on final outcomes while neglecting exploration processes within stateful environments. We propose a novel evaluation framework built upon Roblox Studio, which uniquely incorporates exploratory behavior as an assessment dimension. By decoupling observation and editing tools, designing an eight-tool action space, and implementing executable verification scripts, the framework supports large-scale parallel evaluation. Experimental assessments across thirteen frontier models yield a best single-turn pass rate of 51.7%, demonstrating a significant positive correlation between exploration depth and task success rates. This work fills a critical gap in procedural evaluation metrics for autonomous agents, and the complete task suite has been open-sourced to facilitate further research.
📝 Abstract
We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks. Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks. We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at https://github.com/Roblox/open-game-eval.
Problem

Research questions and friction points this paper is trying to address.

Agentic Programming
Game Engine
Exploration Behavior
Benchmark
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Programming
Stateful Game Engine
Exploration Behavior
Benchmark Framework
Tool Separation
🔎 Similar Papers
E
Eray Turkel
Roblox
M
Mengsha Sun
Roblox
K
Kartik Ayyar
Roblox
S
Sean Dunigan
Roblox
Jack Lu
Jack Lu
New York University
Machine LearningDeep LearningGenerative Modeling
V
Vlad Shcherban
Roblox
H
Hsiang-Shun Shih
Roblox
X
Xin Wang
Roblox
Tiantian Zhang
Tiantian Zhang
Tsinghua University
Reinforcement LearningClusteringData Mining