LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive inference costs and context overflow caused by continuously growing agent memory, noting that conventional summarization methods suffer from high latency and lack theoretical guarantees against information loss. To this end, this work proposes LAM, a framework for efficient lossy memory management that integrates deterministic deduplication, cached prefix preservation, and a performance prediction model. LAM introduces retrieval score error bounds to ensure compression reliability, overlaps compression with inference execution, and establishes a pre-deployment cost estimation mechanism. Experimental results demonstrate that LAM successfully removes 22.47% of observation tokens while retaining 99.984% of critical evidence, thereby achieving a 71.4× to 91.6× end-to-end speedup.
📝 Abstract
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.
Problem

Research questions and friction points this paper is trying to address.

Agent Memory
Context Window
Inference Cost
Information Loss
Summarization Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lossy Agent Memory
Deterministic Deduplication
Retrieval-Score Error Bound
Memory Compaction
Performance Model
🔎 Similar Papers
No similar papers found.