chunk-adaptive retrieval

Designs and implements retrieval and reranking pipelines that adapt retrieval configuration at the individual chunk level: running parallel retrievers across configurations (e.g., modality or granularity), computing chunk-level scores per configuration, reranking chunks by those scores, and selecting a winning configuration for each chunk to produce the final retrieved set.

chunk-adaptiveretrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Drowning in Documents: Consequences of Scaling Reranker Inference

Nov 18, 2024
MJ
Mathew Jacob
🏛️ Databricks

This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.

Assessing reranker effectiveness with modern dense embeddingsEvaluating reranker performance beyond first-stage retrievalIdentifying performance decline in rerankers with document scaling

This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.

component-wise optimizationdense retrievalmulti-stage training

This work addresses the inefficiency of traditional retrieval systems that uniformly apply high-cost reranking models to all queries, incurring unnecessary latency and computational overhead for simple queries. The authors propose a utility-based adaptive reranking framework that dynamically selects reranking strategies according to query complexity, enabling cost-aware query routing. A novel utility function is introduced to guide routing decisions, and the approach leverages BM25 for sparse retrieval, MiniLM-L6-v2 for lightweight dense reranking, and BGE-v2-m3 for heavyweight neural reranking. A trained routing classifier enables multi-tier reranking strategy selection. Compared to applying the full BGE model universally, the proposed method reduces median latency by 1.15× to 53× and average latency by 1.11× to 5.22×, with nDCG@10 varying between –17.5% and +4.0%, demonstrating competitive effectiveness across multiple datasets.

Computational CostInformation RetrievalLatency

In information retrieval (IR) experiments, pipeline-based architectures suffer from redundant computation—e.g., repeated retrieval for multi-ranker comparisons—and design–implementation misalignment due to reliance on intermediate result files. To address these issues, this paper proposes a dual-path caching mechanism. First, we introduce *implicit prefix caching*, a novel technique that automatically identifies and reuses common subcomputations via runtime cache-key derivation and operation-sequence hashing. Second, we design *pyterrier-caching*, a pluggable explicit caching extension supporting persistent storage of intermediate representations and modular integration. Implemented atop PyTerrier, our approach preserves end-to-end semantic integrity while significantly reducing I/O overhead and retrieval latency. Empirical evaluation across realistic IR research workflows demonstrates the method’s effectiveness, generality across diverse experimental configurations, and ease of adoption—requiring minimal code changes and no modification to existing pipelines.

Brittle intermediate result files causing implementation disconnectImproving caching in PyTerrier for efficient pipeline executionRedundant computations in IR pipeline experiments

Latest Papers

What's happening recently
View more

This work addresses the challenge of ineffective reranking in dense retrieval systems under zero-shot scenarios, where supervised signals are absent. The authors propose DART, a novel method that performs lightweight adaptive training at test time to refine reranking. Specifically, DART generates pseudo-labels from top- and bottom-ranked documents in the initial retrieval results and fine-tunes the bilinear scoring matrix via a small number of gradient updates, guided by a confidence-weighted margin loss and a cross-query momentum buffering mechanism. Requiring no additional annotations, DART achieves an average relative improvement of 2.1% in NDCG@10 across six BEIR benchmarks, with less than 10ms added latency per query.

BM25cross-encoderdense retrieval

Fixed-size retrieval struggles to accommodate varying query complexity, often leading to either excessive or insufficient document retrieval. This work proposes ScoreGate, a lightweight adaptive mechanism that dynamically determines the number of retrieved passages by fusing similarity scores from a bi-encoder with reranking scores from a cross-encoder—without requiring additional model invocations. ScoreGate is the first approach to jointly leverage both scoring signals to identify relevant documents previously underestimated due to lexical mismatch, thereby overcoming the limitations of fixed top-K retrieval or single-threshold strategies. Experiments show that on MS MARCO, ScoreGate achieves an MRR@10 of 0.401 while reducing retrieved passages by 35%. In internal evaluations, it attains near-perfect recall with zero false positives, cuts per-query token usage by 34.8%, and introduces only 31ms of latency.

adaptive retrievalfixed-cardinality retrievalquery complexity

This work addresses key limitations in existing re-ranking methods: autoregressive models suffer from high latency and constrained search spaces, while non-autoregressive approaches exhibit weak cross-position coordination and a tendency to generate redundant recommendations. To overcome these challenges, the paper proposes a parallel re-ranking framework guided by optimal transport, which leverages dynamic retrieval indexing to achieve global structural coherence and efficient, duplication-free generation. The core innovations include entropy-regularized optimal transport for conflict-aware training, a prefix-anchored credit assignment mechanism that decomposes list-level rewards into position-level signals, and the integration of continuous latent space mapping with hard-matching inference. Evaluated on large-scale industrial recommendation scenarios, the proposed method significantly outperforms current re-ranking baselines, with both offline and online experiments confirming its effectiveness.

autoregressive modelscross-position coordinationduplicate-free slate

Hot Scholars

SL

Shuochen Liu

University of Science and Technology of China
Large Language Model
HL

Huan Li

ZJU100 Young Professor
AI Data PreparationEfficient AISpatiotemporal Data
NR

Nesar Ramachandra

Computational Scientist, Argonne National Laboratory
CosmologyMachine Learning
KC

Ke Chen

Associate Professor of Computer Science, Zhejiang University
database system
JZ

Jun Zhang

Bosch Security Systems B.V.
Computer VisionMachine LearningImage Processing