PACMI: Provenance-Aware Cascading Memory Invalidation for Long-Term LLM Agents

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in LLM agents' long-term memory where outdated data is difficult to trace for dependencies and prone to losing historical value. To tackle this, we propose a provenance-aware cascading failure framework that constructs a provenance graph with typed dependency edges to represent memories and evidence. A four-state validity model is designed to propagate changes, while a premise checker is introduced to enable cascading failure detection and retrieval optimization. Additionally, a multi-domain diagnostic benchmark is established. Experimental results demonstrate that the proposed framework achieves state-of-the-art final answer accuracy on the benchmark, with the premise checker attaining perfect metrics. It significantly reduces errors caused by stale memories, effectively balancing dependency tracing with historical preservation.
📝 Abstract
LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Long-Term LLM Agents
Memory Invalidation
Dependency Tracking
Outdated Memories
Historical Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance Graph
Cascading Memory Invalidation
Validity Lattice
Long-Term LLM Agents
Stale-Premise Detection
🔎 Similar Papers
No similar papers found.