Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

📅 2026-07-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the individual contributions and interaction effects of executable world models, policy distillation, and verification mechanisms on performance in ARC-AGI-3 tasks. By constructing four nested variants of Codex-based agents and conducting experiments across multiple large language models—including the GPT-5 series—and varying reasoning intensities, the work isolates and quantifies, for the first time, the impact of these components on complex reasoning. Results demonstrate that the full variant integrating all components achieves optimal performance: when using gpt-5.6-sol, it attains 99% RHAE on the public test set and solves all tasks completely, strongly validating the synergistic efficacy of these modules in enhancing agent reasoning capabilities.
📝 Abstract
Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verification; the same executable model with scheduled simplification; and a fixed-interface verification treatment that retains simplification and requires exact reproduction of recorded observations. The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games. Exploratory follow-ups evaluate the textual and verification variants with gpt-5.6-sol at xhigh and max. The most robust result is that every agent variant improves with a stronger model and with greater reasoning effort. Within each model-effort setting, differences among variants are smaller than anticipated, while the effects of individual components vary across settings. Requiring a persistent executable deliverable is not universally beneficial: the textual variant outperforms the flexible-interface executable variant in both gpt-5.5 settings. Simplification improves performance in three of the four model-effort settings, with the weakest setting as the only exception. The complete verification treatment ranks first in all four settings, although it uses substantially more resources. In the gpt-5.6-sol follow-up, the verification variant fully solves every public game at both reasoning efforts, achieves about 99% RHAE, and uses fewer than half the total actions of the human baseline. Because the model postdates these games and held-out performance remains untested, this result should be interpreted as saturation of the public set only.
Problem

Research questions and friction points this paper is trying to address.

ARC-AGI-3
executable world models
simplification
verification
attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable world model
scheduled simplification
replay verification
ARC-AGI-3
ablation study
🔎 Similar Papers
2024-02-08International Conference on Machine LearningCitations: 6