Score
Designs and implements retrieval pipelines that produce an initial candidate pool from document- or chunk-level indexes with fast retrievers (dense or other), then progressively refine and re-rank candidates through cascade stages—e.g., cross-encoder rerankers, cluster-aware negatives, or selective LLM resolvers—to yield a final ranked set. Analyzes and tunes stage-specific chunking, language-specific retrievers, and selection policies to balance precision, candidate coverage, latency, and cost.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This work addresses the inefficiency and limited candidate quality of multi-vector retrieval systems, which typically rely on costly exhaustive per-token retrieval. To overcome these limitations, the authors propose a novel two-stage architecture: in the first stage, a learning-based sparse retriever (LSR) with zero inference overhead replaces conventional token-level collection, substantially reducing query encoding costs; in the second stage, a combination of reranking and early pruning strategies enhances efficiency while preserving retrieval effectiveness. Experimental results demonstrate that the proposed method achieves up to 24× speedup over existing approaches and an overall efficiency gain of 1.8×, all while maintaining comparable or superior retrieval quality.
To address the problem that listwise LLM re-rankers permanently miss highly relevant documents due to insufficient initial retrieval recall, this paper proposes the first adaptive retrieval framework tailored for listwise LLM re-ranking. Departing from the conventional assumption of independent document scoring, our method dynamically generates feedback from LLM re-ranking outputs and leverages it to guide multi-round retrieval in real time; final results are obtained via lightweight fusion of initial-retrieval and feedback-retrieved documents. Extensive experiments across diverse LLM re-rankers, first-stage retrievers, and feedback sources demonstrate improvements of up to 13.23% in nDCG@10 and 28.02% in recall—without increasing LLM inference cost. Our core contribution is the first integration of adaptive retrieval into the listwise LLM re-ranking paradigm, enabling closed-loop, synergistic optimization between retrieval and re-ranking.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.
To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.
This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.
This work addresses the challenge of ineffective reranking in dense retrieval systems under zero-shot scenarios, where supervised signals are absent. The authors propose DART, a novel method that performs lightweight adaptive training at test time to refine reranking. Specifically, DART generates pseudo-labels from top- and bottom-ranked documents in the initial retrieval results and fine-tunes the bilinear scoring matrix via a small number of gradient updates, guided by a confidence-weighted margin loss and a cross-query momentum buffering mechanism. Requiring no additional annotations, DART achieves an average relative improvement of 2.1% in NDCG@10 across six BEIR benchmarks, with less than 10ms added latency per query.
This work addresses the inefficiency of traditional retrieval systems that uniformly apply high-cost reranking models to all queries, incurring unnecessary latency and computational overhead for simple queries. The authors propose a utility-based adaptive reranking framework that dynamically selects reranking strategies according to query complexity, enabling cost-aware query routing. A novel utility function is introduced to guide routing decisions, and the approach leverages BM25 for sparse retrieval, MiniLM-L6-v2 for lightweight dense reranking, and BGE-v2-m3 for heavyweight neural reranking. A trained routing classifier enables multi-tier reranking strategy selection. Compared to applying the full BGE model universally, the proposed method reduces median latency by 1.15× to 53× and average latency by 1.11× to 5.22×, with nDCG@10 varying between –17.5% and +4.0%, demonstrating competitive effectiveness across multiple datasets.