PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

๐Ÿ“… 2026-07-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the limitations of large language model agents in long-horizon tasksโ€”such as ARC-AGI-3โ€”where performance is constrained by perceptual, reasoning, and exploratory capabilities. The authors propose a lightweight, programmatic memory framework that structures and records complete interaction logs, enabling efficient historical retrieval through a code-based agent. This approach preserves full contextual information while overcoming the traditional trade-off between comprehensive memory retention and retrieval efficiency. Evaluated on the full ARC-AGI-3 test set, the method improves pass@1 accuracy by 18.0 percentage points to 76.1% and reduces token consumption by 4.2โ€“5.8ร—. When integrated with the Fable 5 model, it achieves a best@2 accuracy of 97.4%.
๐Ÿ“ Abstract
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of \$1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.
Problem

Research questions and friction points this paper is trying to address.

long-horizon reasoning
large language model agents
context management
continual learning
ARC-AGI-3
Innovation

Methods, ideas, or system contributions that make the work stand out.

programmatic memory
long-horizon reasoning
context management
structured interaction log
coding agents
๐Ÿ”Ž Similar Papers