Human-Inspired Memory Architecture for LLM Agents

📅 2026-05-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of effective long-term memory management in large language model agents during extended interactions. Inspired by human cognition, the authors propose a novel memory architecture that systematically integrates six key mechanisms: sleep-based consolidation, interference-driven forgetting, memory trace maturation, retrieval-induced reconsolidation, entity-centric knowledge graphs, and multi-cue hybrid retrieval. To prevent data leakage, they introduce an unsupervised synthetic calibration method for threshold setting. The framework further incorporates memory deduplication and compression, context budget control, and a streaming multi-level evaluation paradigm. Evaluated on the VSCode dataset, the approach achieves 97.2% memory retention accuracy while reducing storage overhead by 58%. On the LongMemEval benchmark, it matches baseline retrieval accuracy using only 200K tokens of context and improves S-tier preference recall by 13.3 percentage points.
📝 Abstract
Current LLM agents lack principled mechanisms for managing persistent memory across long interaction horizons. We present a biologically-grounded memory architecture comprising six cognitive mechanisms: (1) sleep-phase consolidation, (2) interference-based forgetting, (3) engram maturation, (4) reconsolidation upon retrieval, (5) entity knowledge graphs, and (6) hybrid multi-cue retrieval. Each mechanism addresses a specific failure mode of naive memory accumulation. We introduce a synthetic calibration methodology that derives all pipeline thresholds without benchmark data exposure, eliminating a common source of evaluation leakage. We evaluate on two benchmarks. First, a VSCode issue-tracking dataset (13K issues, 120K events) where deduplication-based consolidation achieves 97.2% retention precision with 58% store reduction (+21.8 pp over baseline). Second, the LongMemEval personal-chat benchmark where we conduct the first streaming M-tier evaluation (475 sessions, ~540K unique turns). At a 200K-token context budget, our pipeline matches raw retrieval accuracy (70.1% vs. 71.2%, overlapping 95% CI) while exposing a tunable accuracy/store-size operating curve. At S-tier scale (50 sessions), dedup-based consolidation yields a +13.3 pp improvement in preference recall.
Problem

Research questions and friction points this paper is trying to address.

persistent memory
LLM agents
long interaction horizons
memory management
memory consolidation
Innovation

Methods, ideas, or system contributions that make the work stand out.

memory consolidation
biologically-inspired architecture
synthetic calibration
hybrid multi-cue retrieval
engram maturation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Doga Kerestecioglu
Microsoft
A
Alexei Robsky
Microsoft
C
Clemens Vasters
Microsoft
A
Anshul Sharma
Microsoft
Y
Yitzhak Kesselman
Microsoft