🤖 AI Summary
This study addresses the loss of critical details caused by premature compression during long-context reasoning in large language models. To mitigate this issue, we propose an Evidence Commitment Memory mechanism coupled with a "deferred compression" strategy. This approach temporarily stores unverified segments in memory and employs a learnable policy network trained via reinforcement learning—incorporating both fine-grained step-level and final-answer rewards—to dynamically determine information retention, thereby ensuring factually grounded reasoning. Under a frozen verifier and chunked processing architecture, the proposed model achieves substantial improvements of 10.4 to 11.4 F1 points over the strongest baseline across long-input scenarios involving 6,400 documents. These results demonstrate that our method significantly enhances the accuracy and reliability of long-text reasoning.
📝 Abstract
Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.