Persistent Context Graphs for Efficient Memory Compaction in LLM Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inefficient memory compression and substantial re-encoding overhead in long-horizon tasks for LLM-based agents by proposing ReCAP. This method introduces a lightweight, persistent context graph that decouples importance assessment from relevance determination. By storing attention-based importance scores and dependency relations, ReCAP dynamically filters critical messages in response to new queries, achieving model-call-free memory compression while effectively circumventing KV cache re-encoding. Experimental results demonstrate that ReCAP reduces compression latency by approximately 95%, decreases historical context length by nearly half, and significantly improves accuracy on code-related tasks.
📝 Abstract
As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.
Problem

Research questions and friction points this paper is trying to address.

Memory Compaction
LLM Agents
Context Window
KV Cache
Interaction History
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory Compaction
Persistent Context Graph
LLM Agents
Attention-derived Importance
KV Cache