Score
Designs, implements, and evaluates methods that combine retrieval outputs from multiple embedding spaces or models into a single ranked result. This includes score- and rank-level fusion strategies (e.g., reciprocal rank fusion), normalization and weighting to balance scores across embeddings, and techniques for merging global and fine-grained semantic signals.
This paper addresses the challenge of fusing multi-source heterogeneous evidence in retrieval-augmented generation (RAG). To tackle the incompatibility of disparate scoring scales across IR models and the limited generalization of single-model retrievers, we propose a hierarchical rank fusion framework: (1) constructing dual retrieval channels—one for labeled and one for unlabeled data; (2) unifying retrieval scores via z-score normalization to harmonize heterogeneous ranking outputs; and (3) integrating cross-source results using a multi-information-source separation–aggregation strategy. Our approach effectively mitigates score incomparability and model-specific bias. Empirical evaluation on fact verification demonstrates consistent superiority over state-of-the-art single-model and single-source baselines in retrieval accuracy, while significantly improving out-of-domain generalization. These results validate the efficacy of co-modeling heterogeneous data and jointly optimizing multiple rankers within a unified fusion architecture.
This work addresses the challenge of effectively fusing multiple heterogeneous retrieval channels under strict latency constraints to optimize business metrics such as user conversion. We propose a channel-aware unified learning-to-rank framework that formulates multi-channel result fusion as a query-dependent multi-objective ranking problem, jointly optimizing for click-through, add-to-cart, and purchase outcomes. The approach explicitly incorporates channel-specific signals and users’ short-term behavioral sequences, and leverages query-adaptive fusion strategies alongside cross-channel interaction modeling to overcome the limitations of conventional fixed-weight fusion methods. Online A/B experiments demonstrate that the system achieves a 2.85% improvement in user conversion rate while maintaining a p95 latency below 50 milliseconds, and has been successfully deployed in the production environment of Target.com.
This study investigates whether retrieval fusion techniques—commonly adopted in real-world retrieval-augmented generation (RAG) systems, such as multi-query retrieval and reciprocal rank fusion—consistently improve end-to-end answer quality under practical deployment constraints. Conducted within an enterprise knowledge-base RAG pipeline, the evaluation is performed under fixed retrieval depth, reranking budget, and latency limits. While retrieval fusion enhances initial recall, it fails to translate into improved Top-k accuracy after subsequent reranking and context truncation; notably, Hit@10 declines from 0.51 to 0.48 and incurs additional latency. These findings challenge the prevailing assumption of the default efficacy of recall-oriented fusion strategies, revealing diminishing returns in production settings where downstream processing and system constraints critically shape overall performance.
This work addresses the challenge of effectively fusing heterogeneous scores—such as vector similarity and graph-based relevance measures like personalized PageRank—in graph-augmented retrieval, where distributional mismatches hinder integration. To resolve this, the authors propose a calibration method based on Percentile Rank (PIT) normalization, which maps disparate scores onto a unified, dimensionless scale while preserving magnitude information and enabling stable alignment. Combined with linear and Boltzmann fusion strategies, the approach significantly improves last-hop retrieval performance in multi-hop question answering. On MuSiQue and 2WikiMultiHopQA benchmarks, it achieves LastHop@5 scores of 76.5% and 53.6%, respectively, substantially outperforming existing baselines.
Existing hybrid search systems lack systematic empirical analysis of trade-offs among lexical and semantic retrieval components—i.e., retrieval paradigms, fusion strategies, and re-ranking methods—leading to complex, suboptimal configurations. Method: We introduce the first benchmark framework tailored for advanced hybrid architectures, conducting systematic evaluation across 11 real-world datasets, covering four retrieval paradigms, their combinations, and re-ranking strategies. Contribution/Results: We identify a “weakest-link” effect in hybrid pipelines and propose a data-driven configuration mapping method. Crucially, we find Tensor-based Re-ranking Fusion (TRF) achieves both high efficiency and strong semantic modeling under low-resource conditions, overcoming traditional fusion bottlenecks. Experiments reveal that hybrid performance is severely constrained by imbalanced path quality; optimal configurations are highly dependent on dataset characteristics and resource constraints. TRF significantly improves the effectiveness–cost trade-off, outperforming state-of-the-art baselines across diverse settings.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This study investigates whether static embeddings retain complementary value in hybrid retrieval systems that combine lexical and Transformer-based approaches. On five Dutch-language tasks from the MTEB-NL benchmark, the authors integrate BM25, Qwen3-Embedding-0.6B, and two multilingual static embeddings using reciprocal rank fusion (RRF), evaluating performance rigorously via ten-fold query-level cross-validation, bootstrap confidence intervals, and sign randomization tests. Results show that fusing BM25 with Qwen significantly improves MRR on four tasks (by up to +0.061), whereas adding static embeddings yields no significant gains and occasionally degrades performance. The optimal fusion weights consistently lie on the BM25–Qwen boundary, supporting a dual-retriever architecture as a robust default and challenging conventional practices that rely on standalone embedding evaluations.
This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.