🤖 AI Summary
This work addresses the challenge of memory management for language agents operating under strict constraints on context window size and reasoning cost, where retaining raw records often exceeds budget limits while compression risks losing critical information. The authors propose a budget-aware mechanism for selecting memory operations, formally decomposing memory utility for the first time to reveal how coverage and replacement effects jointly determine the optimal strategy—whether to retain, merge, abstract, or rewrite. Leveraging an Offline Abstraction Safety (OAS) approach, the method employs pre-generated features and external calibration to efficiently estimate operation utility with minimal overhead, enabling safe and adaptive decisions. Experiments on LongMemEval and LoCoMo benchmarks show that under tight budgets, compression strategies improve accuracy by up to 48%, whereas retention performs better under looser budgets; notably, cross-record abstraction and merging significantly outperform local rewriting.
📝 Abstract
Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their inference costs constrain how much memory can be used at once. Existing systems mainly follow two strategies: memory retention and memory consolidation. Retention keeps raw records and preserves exact details, but relevant evidence may not fit under a tight budget; consolidation compresses and combines records, improving coverage per token but risking the loss of query-critical details. Neither strategy is universally preferable. This raises two central questions: when should consolidation replace retention, and which operator -- Merge, Abstract, or Rewrite -- should be selected? We formalize this decision by decomposing each operator's utility into a coverage effect on evidence omitted by retention and a signed replacement effect on raw evidence that already fits. Their balance explains why the preferred action changes with relative budget pressure. We implement this mechanism with Offline Abstraction-Safety (OAS), a lightweight learner that estimates action utilities from pre-generation features with held-out harm calibration. The public LongMemEval and LoCoMo benchmarks show the same budget-dependent pattern. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets, whereas retention is preferable under loose budgets; LoCoMo replicates this crossover at a smaller budget, consistent with its shorter evidence. On both datasets, cross-note abstraction and merging generally outperform local rewriting when compression is necessary.