Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory

📅 2026-03-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the relative importance of memory writing and retrieval in large language model agents and identifies the primary performance bottleneck. To this end, we propose the first diagnostic framework that disentangles bottlenecks across memory writing, retrieval, and utilization stages. We conduct a 3×3 controlled experiment on the LoCoMo benchmark, evaluating three writing strategies—naive chunking, Mem0-style fact extraction, and MemGPT-style summarization—paired with three retrieval methods: cosine similarity, BM25, and hybrid re-ranking. Results reveal that retrieval methods dominate performance variation, yielding up to a 20-percentage-point accuracy gap, whereas writing strategies exert limited influence. Notably, naive chunking—requiring no LLM calls—matches or surpasses more complex writing approaches, indicating that current systems are primarily constrained by retrieval efficacy rather than writing sophistication.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)

Application Category

Search and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Memory-augmented LLM agents store and retrieve information from prior interactions, yet the relative importance of how memories are written versus how they are retrieved remains unclear. We introduce a diagnostic framework that analyzes how performance differences manifest across write strategies, retrieval methods, and memory utilization behavior, and apply it to a 3x3 study crossing three write strategies (raw chunks, Mem0-style fact extraction, MemGPT-style summarization) with three retrieval methods (cosine, BM25, hybrid reranking). On LoCoMo, retrieval method is the dominant factor: average accuracy spans 20 points across retrieval methods (57.1% to 77.2%) but only 3-8 points across write strategies. Raw chunked storage, which requires zero LLM calls, matches or outperforms expensive lossy alternatives, suggesting that current memory pipelines may discard useful context that downstream retrieval mechanisms fail to compensate for. Failure analysis shows that performance breakdowns most often manifest at the retrieval stage rather than at utilization. We argue that, under current retrieval practices, improving retrieval quality yields larger gains than increasing write-time sophistication. Code is publicly available at https://github.com/boqiny/memory-probe.
Problem

Research questions and friction points this paper is trying to address.

retrieval bottleneck
memory utilization
LLM agent memory
write strategy
retrieval method
Innovation

Methods, ideas, or system contributions that make the work stand out.

memory-augmented LLM agents
retrieval bottleneck
write strategy
diagnostic framework
memory utilization
🔎 Similar Papers
No similar papers found.
B
Boqin Yuan
University of California, San Diego
Y
Yue Su
Carnegie Mellon University
K
Kun Yao
University of North Carolina