🤖 AI Summary
This study addresses the limitation of long-horizon agents constrained by context windows, where textual summaries alone are insufficient to support downstream decision-making. To overcome this, we propose a soft memory token mechanism based on neural memory networks that serves as a residual complement to textual summaries. By drawing an analogy to residual connections along the sequence dimension, this approach preserves critical historical information using minimal input positions, enabling frozen large language models to approximate full-history reasoning. Evaluated on the SummHay benchmark, the proposed method substantially reduces input costs while significantly improving accuracy in multi-turn tasks and effectively mitigating tool-calling errors. These results demonstrate that our approach offers an efficient paradigm for optimizing long-horizon reasoning in large language models.
📝 Abstract
Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.