Score
Design and build systems that index and retrieve documents or document regions using approaches such as inverted and vector indices, feature-based and memory-efficient indexing strategies, and pointer-based referencing so retrieval returns compact references (offsets, pointers, or short metadata) instead of full excerpts. Implement and analyze indexing pipelines, similarity search and lookup performance, storage and memory tradeoffs, and protocols for concise reference returns and rehydration so the referenced content can be fetched or reconstructed while minimizing transmitted output tokens.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
Prior surveys on information retrieval (IR) models conflate architectural design with training methodologies, obscuring the intrinsic evolution of structural innovations in relevance modeling. Method: We systematically trace the architectural progression of IR models—spanning backbone feature extractors and end-to-end relevance modeling—from classical BM25 through CNN/RNN-based rankers to modern BERT dual-encoder and interaction-based architectures, ColBERT, Cross-Encoders, and LLM-based retrievers—explicitly decoupling architecture from training strategy. Contribution/Results: We propose the first longitudinal IR-specific architectural taxonomy, explicitly addressing scalability and adaptability challenges in multimodal, multilingual, and emerging application scenarios. Our framework provides an actionable technology roadmap for industrial system selection and rigorously identifies open research questions and future directions for the academic community.
Multi-vector retrieval models (e.g., ColBERT) balance effectiveness and latency but incur high storage overhead and poor OS paging efficiency due to storing one vector per token. This work proposes the **Fixed-Slot Multi-Vector Encoding Paradigm**, the first approach to decouple document representation from input tokenization: a learnable encoder maps documents of arbitrary length into a fixed-size set of slot vectors, with joint optimization of encoding and ranking modules. Evaluated on MSMARCO and BEIR benchmarks, our method retains over 98% of the original retrieval effectiveness while significantly reducing storage footprint. It also improves disk I/O throughput and cache efficiency, thereby achieving a favorable trade-off among retrieval accuracy, storage economy, and system compatibility.
This work proposes DotVByte, a novel compression algorithm specifically optimized for inner product computation in sparse retrieval systems, addressing the high memory overhead of forward index storage. Building upon and enhancing the StreamVByte integer compression technique, DotVByte significantly reduces memory consumption while preserving retrieval accuracy and maintaining low latency. Experimental results on the MS MARCO dataset demonstrate that DotVByte achieves a superior trade-off among compression ratio, retrieval quality, and computational latency, thereby substantially improving the storage efficiency and practicality of sparse retrieval systems.
This work addresses the scalability challenge in multimodal late-interaction retrieval, where document representations as multiple vectors grow linearly with input length, hindering application to rich media such as images and videos. To mitigate storage and computational overhead under a fixed vector budget, the authors propose a query-agnostic compression approach. Its core innovation is Attention-Guided Clustering (AGC), which leverages attention mechanisms to identify semantically salient regions as cluster centers and aggregates them with learned weights, balancing compression flexibility and retrieval effectiveness. Integrated with sequence resampling, memory tokens, and hierarchical pooling, the method achieves superior retrieval performance over existing parameterized compression strategies across multiple benchmarks—including BEIR, ViDoRe, MSR-VTT, and MultiVENT 2.0—yielding more compact indexes while matching or even surpassing the performance of uncompressed models.
Existing hybrid search systems lack systematic empirical analysis of trade-offs among lexical and semantic retrieval components—i.e., retrieval paradigms, fusion strategies, and re-ranking methods—leading to complex, suboptimal configurations. Method: We introduce the first benchmark framework tailored for advanced hybrid architectures, conducting systematic evaluation across 11 real-world datasets, covering four retrieval paradigms, their combinations, and re-ranking strategies. Contribution/Results: We identify a “weakest-link” effect in hybrid pipelines and propose a data-driven configuration mapping method. Crucially, we find Tensor-based Re-ranking Fusion (TRF) achieves both high efficiency and strong semantic modeling under low-resource conditions, overcoming traditional fusion bottlenecks. Experiments reveal that hybrid performance is severely constrained by imbalanced path quality; optimal configurations are highly dependent on dataset characteristics and resource constraints. TRF significantly improves the effectiveness–cost trade-off, outperforming state-of-the-art baselines across diverse settings.
Scientific document retrieval faces significant challenges due to the scarcity of domain-specific labeled data and the highly specialized nature of technical terminology, which often leads existing methods to suffer from conceptual redundancy or insufficient coverage. To address these limitations, this work proposes an academic concept indexing framework that integrates a structured scholarly taxonomy with large language models to extract and organize key concepts. The framework introduces two novel mechanisms: Concept-Coverage-aware Query Generation (CCQGen) and Concept-Focused Context Expansion (CCExpand), which jointly enhance the retrieval system’s capacity to understand and match scientific semantics. Experimental results demonstrate that the proposed approach substantially improves query quality, concept alignment, and overall retrieval effectiveness, outperforming current state-of-the-art methods on scientific document retrieval benchmarks.
This work addresses the challenge of large index sizes in multi-vector retrieval models, which stem from their long embedding sequences and hinder practical deployment. The study presents the first systematic evaluation of training-free token compression strategies that directly reduce the sequence dimensionality of multi-vector embeddings to lower memory overhead and query latency. By comparing token merging against token pruning, the authors demonstrate that merging achieves a superior trade-off: it substantially shrinks index size while more effectively preserving retrieval performance. These findings establish token merging as a practical and effective solution for enabling efficient multi-vector retrieval without compromising accuracy.
This work addresses the limitations of traditional retrieval-augmented generation (RAG) systems, which rely on server-side computation and suffer from privacy risks, high latency, and substantial storage overhead, while on-device deployment is hindered by resource constraints that compromise both retrieval efficiency and generation quality. To overcome these challenges, the authors propose a unified on-device RAG model that jointly optimizes retrieval and context compression for the first time. By sharing document representations across both tasks, the model enables efficient retrieval and compact context generation within a single architecture, eliminating redundancy from multiple separate models. The approach achieves generation performance comparable to conventional RAG using only approximately one-tenth of the context length, while maintaining embedding storage costs no higher than existing multi-vector retrieval methods, thereby significantly enhancing the practicality and efficiency of on-device RAG.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
This work addresses the limitation of traditional retrieval systems, which treat document representation as a static preprocessing step and thus struggle to adapt to downstream tasks. The authors propose AutoIndex, a novel framework that formulates document representation construction as a learnable program synthesis problem. AutoIndex dynamically generates retrieval-oriented representations by searching over executable transformation programs—such as slicing, augmentation, and normalization—and iteratively refines them using validation feedback. By integrating proxy-guided program search with retrieval quality evaluation, the method enables explicit optimization of document representations. Evaluated on the CRUMB benchmark across all eight tasks, AutoIndex consistently outperforms the full-document BM25 baseline, achieving average improvements of 8.4% in Recall@100 and 8.3% in nDCG@10, with peak gains reaching 30.5% and 43.6%, respectively.
This work addresses the inefficiencies of conventional text chunking in standard retrieval-augmented generation (RAG) systems, which often introduces redundancy, leading to excessive storage costs and degraded retrieval performance. To mitigate this, the authors propose a lightweight pre-index filtering mechanism that integrates semantic similarity, topic coherence, and named entity recognition to selectively prune redundant text chunks prior to indexing. Evaluated through a token-level precision, recall, and intersection-over-union framework, the approach reduces vector index size by 25%–36% while preserving retrieval quality comparable to that of the original system, thereby significantly enhancing the overall efficiency of RAG pipelines.