Score
Design and build retrieval scoring systems that keep candidate (document/passage) representations precomputable and compute token-wise similarities between query tokens and those stored token embeddings at query time, typically aggregating scores with a max-sim or related late-interaction function to produce a final ranking. Analyze and optimize the resulting trade-offs in latency, storage for candidate token embeddings, and aggregation strategy to balance efficiency and effectiveness.
This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.
This work addresses the lack of a systematic understanding of the expressive power of the MaxSim similarity function and its comparison with traditional vector inner products. Through a constructive proof, the authors demonstrate that standard MaxSim can exactly reproduce the inner product of any non-negative k-sparse vectors. They further propose Signed MaxSim to support inner products of real-valued vectors, thereby providing the first theoretical quantification of the representational capacity of late-interaction models. This extension not only overcomes MaxSim’s inherent limitation in handling negative values but also exhibits enhanced modeling capabilities in logical expression and soft OR aggregation. Empirical results show that Signed MaxSim significantly improves out-of-domain generalization on retrieval tasks involving negation: nDCG@10 rises from 0.597 to 1.000 under lexical transfer and surges from 0.008 to 0.788 on purely negated queries.
Fixed-size retrieval struggles to accommodate varying query complexity, often leading to either excessive or insufficient document retrieval. This work proposes ScoreGate, a lightweight adaptive mechanism that dynamically determines the number of retrieved passages by fusing similarity scores from a bi-encoder with reranking scores from a cross-encoder—without requiring additional model invocations. ScoreGate is the first approach to jointly leverage both scoring signals to identify relevant documents previously underestimated due to lexical mismatch, thereby overcoming the limitations of fixed top-K retrieval or single-threshold strategies. Experiments show that on MS MARCO, ScoreGate achieves an MRR@10 of 0.401 while reducing retrieved passages by 35%. In internal evaluations, it attains near-perfect recall with zero false positives, cuts per-query token usage by 34.8%, and introduces only 31ms of latency.
In ColBERT, the Chamfer distance neglects query-term importance, limiting fine-grained matching capability. Method: We propose Importance-Weighted Chamfer Distance (IWCD), which introduces learnable, term-level importance weights solely for query tokens—while preserving precomputed document vectors—by weighting the maximum similarity scores between each query token and all document tokens. These weights are initialized with IDF and optimized via few-shot fine-tuning, requiring no modification to existing vector representations or document encoders. Contribution/Results: IWCD significantly enhances semantic alignment in multi-vector retrieval. Under the BEIR zero-shot setting, it improves Recall@10 by +1.28% on average; with few-shot fine-tuning, the gain rises to +3.66%. This demonstrates the method’s effectiveness, efficiency, and strong generalization across diverse retrieval tasks.
This study addresses the lack of systematic evaluation of the accuracy–cost trade-off across retrieval-augmented generation (RAG) approaches under uniformly scaled corpora. The authors construct a rigorously nested corpus ladder spanning 1K to 512K documents and conduct end-to-end controlled experiments on four RAG paradigms—BM25, dense retrieval, graph-based indexing, and agent search—using a fixed question set and unified evaluation protocol. Their analysis reveals, for the first time within a consistent framework, that BM25 achieves both high accuracy and the lowest computational cost at medium to large scales. While standalone agent-based methods exhibit poor scalability, their hybridization with BM25 attains a consistent accuracy of 69.4% across all corpus sizes, substantially outperforming individual methods. In contrast, graph-based RAG proves difficult to scale due to its prohibitive construction overhead.
This work addresses the GPU memory bottleneck in traditional late-interaction retrieval caused by explicit construction of large similarity tensors during MaxSim computation, which severely limits batch size and scalability. The authors propose Flash-MaxSim, an I/O-aware fused GPU kernel that streams query and document chunks through on-chip SRAM and computes row-wise max reductions in a single pass, enabling exact MaxSim evaluation without materializing intermediate tensors for the first time. The method supports backpropagation, INT8 quantization, and variable-length padding-free scoring, and introduces key innovations including inverse-grid CSR layout and atomic-free gradient reduction. Evaluated on A100/H100 GPUs, Flash-MaxSim achieves 3.9–4.7× faster inference, reduces inference and training memory consumption by 16× and 28× respectively, substantially expands tractable corpus and batch sizes, and maintains 100% top-20 ranking consistency.
This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.
Existing retrieval methods struggle to balance global semantic compression with fine-grained token-level interactions in terminology-intensive tasks, resulting in a pronounced trade-off between indexing cost and retrieval effectiveness. This work proposes a unified multi-granularity retrieval framework that introduces, for the first time, context-dependent variable-length phrases as an intermediate retrieval unit, while preserving uncovered tokens as singletons. The approach integrates importance-guided unit selection with a weighted MaxSim interaction mechanism. Evaluated across 16 scientific, medical, and bilingual retrieval tasks, the phrase-based branch achieves an average improvement of 6.91 macro nDCG@10 over the global branch and approaches token-level performance while reducing document vectors by only 13.7%, substantially enhancing both effectiveness and efficiency in terminology-dense retrieval scenarios.
Existing generative recommender systems suffer from a disconnect between semantic ID (SID) construction and personalized ranking objectives, which limits retrieval performance. This work proposes DIG, a novel framework that unifies ranking and retrieval through the lens of tokenization for the first time: it embeds a tokenizer within a discriminative ranking model and trains the entire system end-to-end, leveraging user-item cross features to guide codebook boundary optimization. Additionally, a user-to-token (u2t) distillation module is introduced to enable efficient inference. By design, the ranking model inherently acquires retrieval capabilities, leading to significant improvements across ranking, retrieval, and joint tasks on three public benchmarks and two industrial datasets.