The 99% Success Paradox: When Near-Perfect Retrieval Equals Random Selection

📅 2026-05-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional retrieval systems often exhibit near-random selectivity at high recall levels, limiting the performance of downstream large language model (LLM) tasks. This work proposes the Bits-over-Random (BoR) metric, which introduces an opportunity-correction mechanism grounded in information theory and models the random baseline using the hypergeometric distribution to quantify the true selectivity of retrieval results. Experiments reveal that BM25 and SPLADE achieve BoR ≈ 0 at K=100, indicating a practical loss of selectivity. In contrast, BoR effectively discriminates system performance across BEIR, SciFact, and MS MARCO benchmarks, approaching theoretical upper bounds and demonstrating broad applicability and practical guidance—particularly in deep retrieval and LLM tool selection scenarios within Retrieval-Augmented Generation (RAG) frameworks.
📝 Abstract
For most of the history of information retrieval (IR), search results were designed for human consumers who could scan, filter, and discard irrelevant information on their own. This shaped retrieval systems to optimize for finding and ranking more relevant documents, but not keeping results clean and minimal, as the human was the final filter. However, LLMs have changed that by lacking this filtering ability. To address this, we introduce Bits-over-Random (BoR), a chance-corrected measure of retrieval selectivity that reveals when high success rates mask random-level performance. We measure selectivity as $BoR = \log_{2}\left(\frac{\mathrm{P}_{obs}}{\mathrm{P}_{rand}}\right)$, where $\mathrm{P}_{rand}$ is the hypergeometric baseline for the chosen success rule (here, coverage: $ \geq1 $ relevant in top-$K$). On the 20 Newsgroups dataset, BM25 and SPLADE both report $>99$% success at $K=100$ (coverage), yet $BoR \approx 0$, indicating random-level selectivity at that depth. When the expected coverage ratio $\left(\frac{K \cdot \bar{R}_{q}}{N}\right)$ exceeds 3-5, the baseline dominates and selectivity collapses. Downstream retrieval-augmented generation (RAG) evaluation confirms this pattern: LLM accuracy can degrade substantially at $K=100$, consistent with the near-zero BoR ceiling. In contrast, BoR remains positive on BEIR/SciFact and on MS MARCO (where 41 systems cluster within 0.2 bits of the theoretical ceiling despite a 13-point recall gap), confirming baseline predictions across sparse and large-scale settings. We further show that the collapse boundary applies to LLM agent tool selection, where small catalog sizes cause selectivity to vanish even with perfect selectors. These findings suggest reporting BoR alongside traditional metrics and reconsidering depth choices when additional retrieval provides negligible selectivity gains while inflating computational costs.
Problem

Research questions and friction points this paper is trying to address.

retrieval selectivity
random-level performance
retrieval-augmented generation
success paradox
information retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bits-over-Random
retrieval selectivity
random baseline
RAG evaluation
information retrieval collapse
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
V
Vyzantinos Repantis
Meta Platforms Inc.
H
Harshvardhan Singh
Meta Platforms Inc.
Tony Joseph
Tony Joseph
University of British Columbia
machine learningcomputer vision
C
Cien Zhang
Meta Platforms Inc.
A
Akash Vishwakarma
Meta Platforms Inc.
S
Svetlana Karslioglu
Meta Platforms Inc.
M
Michael Wyatt Thot
Meta Platforms Inc.
A
Ameya Gawde
Meta Platforms Inc.