Score
Designs and evaluates mechanisms that produce the pool of candidate items for downstream ranking or decision-making, including retrieval models, index queries, nearest-neighbor searches, filters, and combiners. Optimizes recall, coverage, and diversity under constraints such as latency and computational budget by engineering hybrid recall sources, tuning retrieval hyperparameters, crafting candidate filters and deduplication rules, and measuring recall-focused metrics.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
Under information overload, the retrieval stage in recommender systems has long been underappreciated and lacks systematic investigation. This paper presents the first comprehensive survey of retrieval in industrial multi-stage recommendation pipelines, focusing on three core aspects: user-item similarity modeling, efficient indexing mechanisms (e.g., vector search and inverted indices), and training optimization techniques—including dual-tower architectures, contrastive learning, and negative sampling. We introduce a unified evaluation benchmark spanning three public datasets and integrate insights from leading industry practitioners to holistically characterize deployment practices, performance bottlenecks, and engineering challenges. Our work fills a critical gap in the systematic analysis of retrieval and provides both theoretical foundations and practical paradigms for designing accurate, efficient, and production-ready retrieval components within cascaded recommendation systems.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
Existing hybrid search systems lack systematic empirical analysis of trade-offs among lexical and semantic retrieval components—i.e., retrieval paradigms, fusion strategies, and re-ranking methods—leading to complex, suboptimal configurations. Method: We introduce the first benchmark framework tailored for advanced hybrid architectures, conducting systematic evaluation across 11 real-world datasets, covering four retrieval paradigms, their combinations, and re-ranking strategies. Contribution/Results: We identify a “weakest-link” effect in hybrid pipelines and propose a data-driven configuration mapping method. Crucially, we find Tensor-based Re-ranking Fusion (TRF) achieves both high efficiency and strong semantic modeling under low-resource conditions, overcoming traditional fusion bottlenecks. Experiments reveal that hybrid performance is severely constrained by imbalanced path quality; optimal configurations are highly dependent on dataset characteristics and resource constraints. TRF significantly improves the effectiveness–cost trade-off, outperforming state-of-the-art baselines across diverse settings.
This work addresses the challenge that existing retrieval methods lack theoretical guarantees when scaling up and struggle to balance relevance and semantic diversity. The authors formulate diversity-aware retrieval as a cardinality-constrained binary quadratic programming problem, introducing an interpretable parameter to explicitly trade off between relevance and diversity. They propose, for the first time, a scalable optimization framework with provable convergence guarantees, which combines a non-convex tight continuous relaxation with the Frank-Wolfe algorithm to enable efficient solution. Experimental results demonstrate that the proposed method consistently outperforms existing baselines across the relevance-diversity Pareto frontier while achieving substantial gains in computational efficiency.
This work addresses the challenge of jointly optimizing multiple retrieval models—such as BM25 and large language models (LLMs)—within composite retrieval systems. We propose the first end-to-end learnable paradigm for composite retrieval, which simultaneously optimizes *where* to invoke each model (e.g., deploying LLMs for relative relevance judgment rather than conventional top-K re-ranking) and *how* to fuse their predictions, jointly maximizing both retrieval effectiveness (NDCG) and computational efficiency. Our method integrates self-supervised training, multi-stage scheduling, and a learnable weighted fusion mechanism, enabling optimization without human annotations. Evaluated on multiple benchmarks, our approach achieves a 3.2% absolute improvement in NDCG@10 over cascaded baselines under identical computational budgets—significantly surpassing the limitations of traditional cascade-based re-ranking. This establishes a new, efficient, and scalable paradigm for hybrid retrieval systems.
This work addresses the significant degradation in retrieval efficiency caused by fragmented connectivity in traditional graph indexes when handling low-selectivity filtered queries, where qualifying vectors are sparse. To overcome this limitation, the authors propose Curator, a partitioned dual-index architecture based on shared clustering trees that constructs dedicated sub-indexes for different labels, thereby maintaining high search efficiency while reducing memory overhead. Curator introduces an adaptive partitioning mechanism that supports incremental updates and enables on-the-fly temporary index construction during query execution, effectively accommodating complex predicate filtering. Experimental results demonstrate that, when integrated with state-of-the-art graph indexes, Curator reduces query latency by up to 20.9× for low-selectivity queries, with only a 5.5% increase in index construction time and a 4.3% increase in memory usage.
Traditional retrieval systems often exhibit near-random selectivity at high recall levels, limiting the performance of downstream large language model (LLM) tasks. This work proposes the Bits-over-Random (BoR) metric, which introduces an opportunity-correction mechanism grounded in information theory and models the random baseline using the hypergeometric distribution to quantify the true selectivity of retrieval results. Experiments reveal that BM25 and SPLADE achieve BoR ≈ 0 at K=100, indicating a practical loss of selectivity. In contrast, BoR effectively discriminates system performance across BEIR, SciFact, and MS MARCO benchmarks, approaching theoretical upper bounds and demonstrating broad applicability and practical guidance—particularly in deep retrieval and LLM tool selection scenarios within Retrieval-Augmented Generation (RAG) frameworks.
This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.
This work addresses the challenge of expansion bias in large-scale retrieval systems, which disproportionately affects fresh and long-tail content, leading to uneven model performance gains. To mitigate this, the authors propose MESH, a unified retrieval expansion framework that incorporates structural inductive biases through a modular neural architecture, feature space partitioning, and a gated bias-correction mechanism. This design preserves gradient pathways for sparse items, effectively decoupling interference between high-frequency and sparse signals while enabling asynchronous inference to enhance throughput. Evaluated on Pinterest’s billion-scale recommendation system, MESH achieves a 5.5% increase in re-pin rate for fresh content, a 55% improvement in funnel efficiency, a 0.46% gain in user retention, and a 2.87× boost in system throughput.
This work addresses the challenge that existing hybrid vector-attribute query methods struggle to simultaneously achieve efficiency, generality, and graph connectivity under low-selectivity conditions. The authors propose an efficient, filter-oblivious approximate nearest neighbor search framework that dynamically optimizes query paths by unifying selectivity estimation with search execution. Key innovations include a selectivity-aware exclusion distance mechanism to reshape distance distributions, a selectivity-driven dynamic routing strategy, and an inline filtering algorithm built upon HNSW. Experimental results demonstrate that, at Recall@10 = 95%, the proposed method achieves 1.3–5× higher queries per second (QPS) than state-of-the-art approaches on real-world datasets, while matching the performance of specialized solutions under specific filtering constraints.