Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the incoherent world understanding and belief traps encountered by large language model agents during long-horizon tasks, which stem from deficiencies in memory organization. To overcome these challenges, we propose the PoS reasoning framework, which innovatively constructs an explicit belief state—comprising world estimates and unresolved requirements—as the decision context. Furthermore, the framework employs consistency verification to detect belief traps and introduces customized recovery strategies tailored to specific trap patterns and requirement types. Extensive experiments across four benchmarks demonstrate that the proposed method achieves state-of-the-art performance. The effectiveness of its core components is thoroughly validated, and the framework exhibits strong robustness against context growth.
📝 Abstract
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
belief states
long-horizon tasks
belief trapping
context management
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explicit Belief States
Belief Trapping
Long-Horizon Agents
Inference-time Framework
Context Management