AGI Maze as a Benchmark Framework for World-Modeling Agents

📅 2026-07-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) struggle to construct persistent, actionable internal representations of the external world, particularly in partially observable, stateful environments that demand memory and structured reasoning. This work proposes AGI Maze—a lightweight grid-based maze benchmark that, for the first time, offers a standardized evaluation setting focused exclusively on world modeling capabilities without requiring high-dimensional perceptual inputs. The framework employs multi-difficulty maze tasks coupled with a working memory mechanism that retains message history, enabling assessment of whether agents can learn and leverage internal environmental representations. Experimental results demonstrate that even when augmented with working memory, current LLMs fail to reliably build accurate internal models of maze structures within the number of steps humans find trivial, revealing fundamental limitations in their runtime world modeling abilities.
📝 Abstract
Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world. Many tasks that look like "reasoning" in text become substantially harder once the environment is partially observable, stateful, and requires memory and structured hypotheses about hidden state. AGI Maze is a lightweight framework for building such environments without requiring high-dimensional sensory inputs. It provides a family of grid-based maze tasks with a clean API and multiple difficulty regimes. The goal is to create benchmarks where agents must learn and use world state representations, not just infer a local rule over readily provided observations. We provide an initial evaluation of several vanilla LLMs on simple mazes showing that they fail to represent mazes internally at LLM inference time. We also introduce a baseline agent, which is allowed to use its message history as a working memory to construct descriptions of observations at agentic runtime. Although this can improve performance, it is still insufficient for an LLM agent to reliably solve even small mazes within a step budget that is more than enough for humans.
Problem

Research questions and friction points this paper is trying to address.

world modeling
partially observable environments
memory
state representation
agent benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

world modeling
partially observable environments
AGI Maze
memory-augmented agents
LLM benchmarking