Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in Retrieval-Augmented Generation (RAG), where retrieved passages, though relevant, often lack key facts, causing models to hallucinate answers despite insufficient evidence. To mitigate this, we propose RINSE, the first framework that determines evidence sufficiency without requiring generation. By integrating three signals—question coverage, answer presence, and small-model judgment—RINSE effectively circumvents interference from parametric knowledge. Furthermore, we construct a benchmark controlling for surface features to resolve misclassification issues. Experiments demonstrate that RINSE achieves an AUC of 0.837 across six datasets, outperforming existing methods and frontier API-based models. With a single-GPU inference latency of merely 36.5ms, RINSE enables efficient and lightweight evidence sufficiency detection.
📝 Abstract
Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Evidence Sufficiency
Evidence Gap Detection
Abstention
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Evidence Sufficiency
RINSE
Abstention Detection
Paired Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Suting Chen
Northwestern University
P
Peichun Hua
The Chinese University of Hong Kong, Shenzhen
Yunming Xiao
Yunming Xiao
The Chinese University of Hong Kong, Shenzhen
Computer NetworksSecurity and PrivacyCloud Infrastructure