multi-stage retrieval

Designs and implements retrieval pipelines that produce an initial candidate pool from document- or chunk-level indexes with fast retrievers (dense or other), then progressively refine and re-rank candidates through cascade stages—e.g., cross-encoder rerankers, cluster-aware negatives, or selective LLM resolvers—to yield a final ranked set. Analyzes and tunes stage-specific chunking, language-specific retrievers, and selection policies to balance precision, candidate coverage, latency, and cost.

multi-stageretrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiency and limited candidate quality of multi-vector retrieval systems, which typically rely on costly exhaustive per-token retrieval. To overcome these limitations, the authors propose a novel two-stage architecture: in the first stage, a learning-based sparse retriever (LSR) with zero inference overhead replaces conventional token-level collection, substantially reducing query encoding costs; in the second stage, a combination of reranking and early pruning strategies enhances efficiency while preserving retrieval effectiveness. Experimental results demonstrate that the proposed method achieves up to 24× speedup over existing approaches and an overall efficiency gain of 1.8×, all while maintaining comparable or superior retrieval quality.

first-stage retrievergather-and-refinemultivector retrieval

Guiding Retrieval using LLM-based Listwise Rankers

Jan 15, 2025
MR
M. Rathee
🏛️ L3S Research Center | University of Glasgow | Delft University of Technology

To address the problem that listwise LLM re-rankers permanently miss highly relevant documents due to insufficient initial retrieval recall, this paper proposes the first adaptive retrieval framework tailored for listwise LLM re-ranking. Departing from the conventional assumption of independent document scoring, our method dynamically generates feedback from LLM re-ranking outputs and leverages it to guide multi-round retrieval in real time; final results are obtained via lightweight fusion of initial-retrieval and feedback-retrieved documents. Extensive experiments across diverse LLM re-rankers, first-stage retrievers, and feedback sources demonstrate improvements of up to 13.23% in nDCG@10 and 28.02% in recall—without increasing LLM inference cost. Our core contribution is the first integration of adaptive retrieval into the listwise LLM re-ranking paradigm, enabling closed-loop, synergistic optimization between retrieval and re-ranking.

Adaptive Retrieval TechniquesLarge Language ModelsSearch Result Relevance

This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.

component-wise optimizationdense retrievalmulti-stage training

Drowning in Documents: Consequences of Scaling Reranker Inference

Nov 18, 2024
MJ
Mathew Jacob
🏛️ Databricks

This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.

Assessing reranker effectiveness with modern dense embeddingsEvaluating reranker performance beyond first-stage retrievalIdentifying performance decline in rerankers with document scaling

RankLLM: A Python Package for Reranking with LLMs

May 25, 2025
SS
Sahel Sharifymoghaddam
🏛️ University of Waterloo

To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.

Addresses reliability concerns in LLM APIs and MoE modelsDevelops modular Python package for LLM-based document rerankingEnables quick reproduction of results for research and applications

Latest Papers

What's happening recently
View more

This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.

cascading pipelinesexperiment efficiencyinformation retrieval

This work addresses the challenge of ineffective reranking in dense retrieval systems under zero-shot scenarios, where supervised signals are absent. The authors propose DART, a novel method that performs lightweight adaptive training at test time to refine reranking. Specifically, DART generates pseudo-labels from top- and bottom-ranked documents in the initial retrieval results and fine-tunes the bilinear scoring matrix via a small number of gradient updates, guided by a confidence-weighted margin loss and a cross-query momentum buffering mechanism. Requiring no additional annotations, DART achieves an average relative improvement of 2.1% in NDCG@10 across six BEIR benchmarks, with less than 10ms added latency per query.

BM25cross-encoderdense retrieval

This work addresses the inefficiency of traditional retrieval systems that uniformly apply high-cost reranking models to all queries, incurring unnecessary latency and computational overhead for simple queries. The authors propose a utility-based adaptive reranking framework that dynamically selects reranking strategies according to query complexity, enabling cost-aware query routing. A novel utility function is introduced to guide routing decisions, and the approach leverages BM25 for sparse retrieval, MiniLM-L6-v2 for lightweight dense reranking, and BGE-v2-m3 for heavyweight neural reranking. A trained routing classifier enables multi-tier reranking strategy selection. Compared to applying the full BGE model universally, the proposed method reduces median latency by 1.15× to 53× and average latency by 1.11× to 5.22×, with nDCG@10 varying between –17.5% and +4.0%, demonstrating competitive effectiveness across multiple datasets.

Computational CostInformation RetrievalLatency

Hot Scholars

MZ

Matei Zaharia

UC Berkeley and Databricks
Distributed SystemsMachine LearningDatabasesSecurity
CP

Christopher Potts

Professor of Linguistics and, by courtesy, of Computer Science
LinguisticsComputational LinguisticsSemanticsPragmatics
JH

Jiawei Han

Abel Bliss Professor of Computer Science, University of Illinois
data miningdatabase systemsdata warehousinginformation networks
TY

Tong Yang

Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookup
DD

Devdatt Dubhashi

Professor Chalmers
AlgorithmsComputational BiologyComputational Modelling