🤖 AI Summary
This study addresses the problem of reasoning over-reliance in agent memory caused by evidence overlap. To mitigate this issue, we propose MEMTRIM, a framework that employs an "index-at-write, control-at-read" memory pruning mechanism to precisely trim redundant memory evidence and alleviate the model's excessive dependence on local information. Functioning as a training-free, plug-and-play module, MEMTRIM effectively distinguishes and preserves critical memory information. Experimental results demonstrate that this framework significantly reduces over-reliance across various large language models and memory architectures while fully preserving downstream task utility, exhibiting strong generalization capability and practical value.
📝 Abstract
Agentic memory allows LLM agents to reuse past experience, yet retrieved memories can also distort inference even when they are benign, correctly stored, and appropriately retrieved. We study this failure mode, which we call memory over-reliance. Across benchmarks and memory architectures, we find that memory is useful when past experience transfers to the current task, but can become misleading when only part of the evidence transfers. Failures are strongest under partial query-memory overlap, a pattern further confirmed by controlled experiments thatvary the amount of overlapping evidence. Motivated by this finding, we propose MEMTRIM, a plug-and-play framework that indexes memory evidence at write time and controls its reuse at read time. MEMTRIM removes repeated or conflicting evidence while preserving useful memory-specific information, requires no retraining, and applies to both embedding-based and structured memory systems.Experiments show that MEMTRIM reduces memory overreliance while preserving the benefits of useful memory across models and memory settings.