A Systematic Multi-Domain Evaluation of Document Retrievers

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmented evaluation landscape of document retrievers, which impedes comprehensive comparisons. Adopting an out-of-the-box paradigm under a unified computational budget, we conduct a large-scale empirical assessment of 33 retrievers—spanning sparse, dense, and augmented architectures—across seven benchmark datasets. Through standardized configurations and multi-dimensional quantitative analysis, we systematically characterize the trade-offs among retrieval effectiveness, latency, and failure modes across different model families. Our findings reveal that NV-Embed-v2 achieves superior performance albeit with higher latency, SPLADE-v3 offers an optimal balance between low latency and high effectiveness, and GritLM demonstrates the strongest instruction-following capability. Furthermore, this work delineates the potential optimization space for each architectural paradigm, providing actionable insights for future retriever development and deployment.
📝 Abstract
Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers' failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.
Problem

Research questions and friction points this paper is trying to address.

Document Retrieval
Multi-Domain Evaluation
Information Retrieval
Retriever Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Document Retrieval
Multi-Domain Evaluation
Off-the-shelf Benchmarking
Failure Point Analysis
Information Retrieval
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.