🤖 AI Summary
Existing long-horizon large language model agents rely heavily on semantic similarity or full-history retrieval for memory management, which often fails to effectively filter out memories that are relevant yet unhelpful, outdated, or harmful, leading to unstable outputs. This work proposes Causal Memory Intervention (CMI), a novel approach that introduces causal reasoning into long-term memory selection for the first time. By applying controlled interventions to quantify the causal effect of candidate memories on model outputs, CMI dynamically retains beneficial memories while suppressing interfering ones. We construct a structured memory repository and introduce Causal-LoCoMo, the first benchmark with causal annotations for evaluating memory systems. Experiments demonstrate that CMI significantly outperforms baseline methods—including vector retrieval, graph-based memory, and reflection mechanisms—on this benchmark, achieving a superior trade-off between response quality and robustness against misleading information.
📝 Abstract
Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using semantic similarity or broad history inclusion, treating retrieved memories as uniformly useful. This assumption is fragile because memories may be topically related while remaining irrelevant, stale, or misleading. We propose Causal Memory Intervention (CMI), a causal memory-selection technique that estimates how candidate memories affect the model's answer under controlled interventions, selecting memories that improve task performance while suppressing unstable, irrelevant, or harmful ones. To evaluate this setting, we introduce Causal-LoCoMo, a causally annotated benchmark derived from long conversational data, where each example contains a user request, a structured memory bank, useful memories, irrelevant distractors, and synthetic harmful memories. We compare CMI against vector, graph, reflection, summary, full-history, and no-memory baselines. Results show that CMI achieves a stronger balance between answer quality and robustness to misleading memory, suggesting that reliable long-term memory requires selecting context based on causal usefulness rather than relevance alone. The full framework, benchmark construction code, and experimental pipeline are available at https://github.com/Saksham4796/causal-memory-intervention.