🤖 AI Summary
This study addresses a critical limitation in Retrieval-Augmented Generation (RAG), where retrieved passages, though relevant, often lack key facts, causing models to hallucinate answers despite insufficient evidence. To mitigate this, we propose RINSE, the first framework that determines evidence sufficiency without requiring generation. By integrating three signals—question coverage, answer presence, and small-model judgment—RINSE effectively circumvents interference from parametric knowledge. Furthermore, we construct a benchmark controlling for surface features to resolve misclassification issues. Experiments demonstrate that RINSE achieves an AUC of 0.837 across six datasets, outperforming existing methods and frontier API-based models. With a single-GPU inference latency of merely 36.5ms, RINSE enables efficient and lightweight evidence sufficiency detection.
📝 Abstract
Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.