Before Answering: Evidence Sufficiency under Size-Matched Memory Construction

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical flaw in existing evidence insufficiency detection benchmarks, where samples constructed via paragraph deletion allow memory size to leak labels, causing models to rely on shortcuts rather than semantic understanding. To overcome this, we propose a size-matched memory construction paradigm that theoretically eliminates such shortcuts. Furthermore, we develop MemSafe, an estimator that models query-memory interactions via a cross-encoder and aggregates features using a Set Transformer to precisely assess evidence sufficiency. Experiments demonstrate that our approach achieves AUROC scores of 0.968 and 0.983 on MuSiQue and HotpotQA, respectively. When employed as a gating mechanism for a 7B-parameter reader, it reduces the error rate from 0.85 to 0.631, significantly enhancing the robustness of missing evidence detection in multi-hop question answering.
📝 Abstract
Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of $0.979$ for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches $0.968$ and $0.983$ AUROC on MuSiQue and HotpotQA, $0.26$ to $0.39$ above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes $41\%$ of the MuSiQue gap between the lexical baseline and MemSafe. At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only $0.639$ AUROC on SQuAD~2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from $0.850$ to $0.631$ at $5\%$ coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at $10\%$ coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.
Problem

Research questions and friction points this paper is trying to address.

evidence sufficiency
memory construction
shortcut learning
multi-hop question answering
unanswerable questions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence Sufficiency
Size-Matched Construction
Cross-Encoder
Set Transformer
MemSafe
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Joyanta Jyoti Mondal
Department of Computer and Information Sciences, University of Delaware, USA
M
Md. Shifatul Ahsan Apurba
Department of Biomedical Informatics and Data Science, University of Alabama at Birmingham, USA
M
Mridul Banik
Luddy School of Informatics, Computing, and Engineering, Indiana University Indianapolis, USA
M
Md Masud Al Mahmud
Department of Computer Science and Engineering, BRAC University, Bangladesh
Ibne Farabi Shihab
Ibne Farabi Shihab
Iowa State University
Deep LearningroboticsLarge Language Model