Exact Adaptive Hybrid Retrieval Without Fixed Top-L Cutoffs

📅 2026-08-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional hybrid retrieval relies on fixed top-L truncation to fuse dense and sparse results, which often introduces ranking bias and struggles to adapt to dynamic changes in queries and corpora. This work proposes Exact Adaptive Hybrid Retrieval (EAHR), the first method to achieve exact adaptive fusion—identical to full-list weighted reciprocal rank fusion (RRF)—without predefining a top-L cutoff. EAHR dynamically adjusts per-channel retrieval depth and continues retrieval only when it impacts the final top-K results. By integrating per-vector scalar quantization and posting block-max techniques, EAHR enables recoverable exact ranking and incorporates a fusion-bound pruning strategy. Experiments demonstrate that EAHR exactly reproduces full top-20 results across five datasets and 150 query–snapshot combinations, achieving 23.35–30.28× speedup over exhaustive batch processing on average.
📝 Abstract
Modern retrieval-augmented generation (RAG) systems often fuse fixed Top-$L$ results from dense and sparse retrievers, treating later contributions as zero. The cutoff therefore determines both the ranking and its execution cost. Yet truncated fusion is not generally equivalent to complete-list fusion: unread cross-list ranks can change Top-$K$ membership or order even when the observed candidates contain every item in the complete-list Top-$K$. Because channel rankings vary across queries and corpus updates, a depth selected from historical queries may not transfer reliably. We propose Exact Adaptive Hybrid Retrieval (EAHR), which fixes the ordered Top-$K$ defined by complete-list weighted RRF as the retrieval target and treats channel depth as request-specific execution state. Per-Vector Scalar Quantization (PVS) and Posting Block-Max (PBM) produce resumable exact dense and sparse rankings. Fusion bounds unread contributions and requests further ranks only while they can change the Top-$K$. Every successful request therefore matches complete-list fusion without a preset Top-$L$; otherwise, execution continues safely to list exhaustion. Across five test collections and five temporal corpus snapshots, complete-list weighted RRF remained competitive, whereas fixed depths selected from historical queries did not transfer reliably. EAHR reproduced the complete-list ordered Top-20 in all 150 query-snapshot combinations. Under a warm-cache, interleaved, order-balanced protocol, the paired geometric-mean latency ratios of exhaustive batch execution to EAHR were 23.35 on TREC-DL 2019 and 30.28 on TREC-DL 2020. Anti-correlated rankings exhausted both lists, and some difficult queries were slower with EAHR. EAHR does not guarantee a speedup for every request; it fixes the exact result while adapting execution depth to the current rankings.
Problem

Research questions and friction points this paper is trying to address.

retrieval-augmented generation
hybrid retrieval
top-K ranking
fixed cutoff
ranking fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Exact Adaptive Hybrid Retrieval
Dynamic Retrieval Depth
Weighted RRF
Per-Vector Scalar Quantization
Posting Block-Max
🔎 Similar Papers