Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in long-term memory question answering where retrieved records are individually relevant yet collectively insufficient as evidence. Drawing upon the concept of sufficiency from legal evidence theory, this work reformulates memory retrieval as constructing a sufficient evidence set. A budgeted flat reconstruction method is proposed to ensure coverage of complementary facts, accompanied by a two-stage algorithm that optimizes retrieval performance under fixed memory constraints. Furthermore, memory selection incorporates formal concept analysis, integrating blind large language model evaluation with multi-view textual search techniques. Evaluated on the LongMemEval-S benchmark, the proposed approach achieves an accuracy of 82.2% and a Turn Hit rate of 91.4%, significantly outperforming existing systems.
📝 Abstract
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.
Problem

Research questions and friction points this paper is trying to address.

long-term memory QA
evidence sufficiency
memory retrieval
LLM agents
multi-session reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Term Memory QA
Budgeted Flat Reconstruction
Formal Concept Analysis
Evidence Sufficiency
Memory Retrieval
🔎 Similar Papers
No similar papers found.