🤖 AI Summary
This study addresses the challenge of ambiguous world models and hindered cognitive revision in agents operating under limited observations within unfamiliar environments. We propose a recursive self-improvement paradigm grounded in explicit semantic memory and code verification. By freezing a large language model to maintain a revisable natural-language rulebook compiled into executable code for predictive planning, our method constructs an observation-reflection-verification closed loop. Prediction errors drive rule revision validated through precise unit-test replay, enabling multi-model parallel exploration guided by shared evidence. Empirically, the approach clears all public games in ARC-AGI-3, achieving 100.0 human-normalized action efficiency while requiring only 44% of the steps, and attains a zero-loss record on Atari Pong.
📝 Abstract
Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.