Score
Designs and builds dense vector retrieval systems that encode queries and passages into dense embeddings, index them with exact or approximate structures (e.g., HNSW), and perform similarity-based top-k search with reranking and fusion strategies to produce ranked passages. Trains and analyzes dense retrieval models—including mining hard negatives and optimizing for scalability, low-latency or on-device constraints—and integrates retrieval with generation and retrieval-augmented evaluation or modeling pipelines.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This paper identifies a fundamental theoretical limitation of the vector embedding retrieval paradigm: even for simple queries, the embedding dimension strictly bounds the number of distinguishable document subsets—a bottleneck intrinsic to the representation and unresolvable via larger models or improved training data. Method: Drawing on statistical learning theory, the authors formally derive an upper bound on the expressive capacity of single-vector embeddings, then construct LIMIT, the first benchmark explicitly designed to probe this theoretical limit; they further propose a minimal empirical framework with k=2 and fully parameterized embeddings for boundary testing. Results: Experiments show that state-of-the-art embedding models underperform significantly relative to the derived theoretical ceiling on LIMIT, directly challenging the prevailing assumption that scaling model size alone suffices to overcome retrieval limitations—and thereby establishing a structural capacity ceiling for the single-vector paradigm.
This study investigates the impact of embedding dimensionality on performance in dense retrieval and its limitations as task complexity increases. Through systematic experiments across models of varying scales, the work presents the first empirical evidence that retrieval performance follows a power-law relationship with embedding dimensionality. Building on this observation, the authors propose predictable scaling laws based solely on dimensionality or jointly on model size. Using dense retrieval architectures, approximate nearest neighbor search, and large-scale comparative evaluations, they demonstrate that in task-aligned scenarios, performance improves with higher dimensionality—albeit with diminishing returns—whereas in misaligned tasks, excessive dimensions degrade performance. These findings offer both theoretical grounding and practical guidance for selecting optimal embedding dimensions in efficient retrieval systems.
This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.
In vector retrieval, inconsistent embedding quality causes significant query-level performance fluctuations, while existing methods lack lightweight, training-free capabilities for per-query performance prediction. To address this, we propose Q-Robust—a fine-tuning-free, lightweight framework that jointly models geometric robustness (quantifying local stability) and neighborhood density (capturing semantic compactness) in the embedding space to accurately predict retrieval effectiveness for individual queries. Our analysis reveals systematic patterns in query-specific embedding quality distributions, enabling dynamic, adaptive retrieval strategies. Evaluated on four standard benchmarks, Q-Robust achieves an average 9.4±1.2% improvement in Recall@10 over strong baselines, with prediction overhead accounting for less than 5% of retrieval latency—demonstrating superior accuracy, efficiency, and practicality.
This work addresses the inefficiency and limited candidate quality of multi-vector retrieval systems, which typically rely on costly exhaustive per-token retrieval. To overcome these limitations, the authors propose a novel two-stage architecture: in the first stage, a learning-based sparse retriever (LSR) with zero inference overhead replaces conventional token-level collection, substantially reducing query encoding costs; in the second stage, a combination of reranking and early pruning strategies enhances efficiency while preserving retrieval effectiveness. Experimental results demonstrate that the proposed method achieves up to 24× speedup over existing approaches and an overall efficiency gain of 1.8×, all while maintaining comparable or superior retrieval quality.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
This work addresses the challenge of ineffective reranking in dense retrieval systems under zero-shot scenarios, where supervised signals are absent. The authors propose DART, a novel method that performs lightweight adaptive training at test time to refine reranking. Specifically, DART generates pseudo-labels from top- and bottom-ranked documents in the initial retrieval results and fine-tunes the bilinear scoring matrix via a small number of gradient updates, guided by a confidence-weighted margin loss and a cross-query momentum buffering mechanism. Requiring no additional annotations, DART achieves an average relative improvement of 2.1% in NDCG@10 across six BEIR benchmarks, with less than 10ms added latency per query.
This study addresses the theoretical limitations of traditional dense retrieval and the vulnerability of generative retrieval under document identifier ambiguity. For the first time, it systematically evaluates the potential of generative retrieval on the synthetic dataset LIMIT, employing SEAL and MINDER models with BM25 and dense retrieval as baselines. Results show that on the original LIMIT dataset, generative approaches achieve R@2 scores of 0.92–0.99, substantially outperforming dense retrieval (<0.03) and BM25 (0.86). However, when hard negatives are introduced, performance sharply drops to 0.51, revealing a critical bottleneck: the decoding mechanism struggles to generate unique identifiers reliably. This work not only confirms the superiority of generative retrieval under ideal conditions but also, through error analysis, identifies identifier ambiguity as a key challenge limiting its robustness.