Learning to Retrieve Missing Evidence for Long-Term Memory QA

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses retrieval failures in long-term memory question answering caused by scattered evidence and missing retrieval cues. We propose MERA, a framework that decouples global memory from evidence states, leveraging verified evidence to guide multi-round iterative retrieval for recovering missing information. Furthermore, this work introduces the first reinforcement learning-based lightweight retrieval planner (0.6B parameters), which optimizes conditional retrieval decisions via reward mechanisms and collaborates with large language models for evidence processing and answer generation. Evaluated on the LoCoMo and LongMemEval-S datasets, MERA achieves accuracies of 77.40% and 71.29%, respectively, surpassing 30B-parameter baseline models while improving cumulative evidence recall to 80.5%.
📝 Abstract
Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.
Problem

Research questions and friction points this paper is trying to address.

Long-term memory QA
Evidence retrieval
Missing evidence
Multi-hop reasoning
Retrieval augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Missing-Evidence Retrieval
Long-Term Memory QA
Reinforcement Learning
Retrieval Augmentation
Lightweight Planner
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.