Score
Designs, builds, and evaluates systems that locate and return relevant items from large collections in response to user queries. This work includes constructing indexing and retrieval pipelines, ranking and relevance models, query processing and expansion, relevance feedback mechanisms, and empirical evaluation with appropriate metrics and test collections.
Prior surveys on information retrieval (IR) models conflate architectural design with training methodologies, obscuring the intrinsic evolution of structural innovations in relevance modeling. Method: We systematically trace the architectural progression of IR models—spanning backbone feature extractors and end-to-end relevance modeling—from classical BM25 through CNN/RNN-based rankers to modern BERT dual-encoder and interaction-based architectures, ColBERT, Cross-Encoders, and LLM-based retrievers—explicitly decoupling architecture from training strategy. Contribution/Results: We propose the first longitudinal IR-specific architectural taxonomy, explicitly addressing scalability and adaptability challenges in multimodal, multilingual, and emerging application scenarios. Our framework provides an actionable technology roadmap for industrial system selection and rigorously identifies open research questions and future directions for the academic community.
Existing retrieval evaluation benchmarks predominantly rely on simple, single-point queries, failing to reflect model capabilities under realistic, complex retrieval scenarios involving multiple constraints and intents. Method: We introduce ComplexRetrieval-Bench—the first systematic, diverse, and realistic benchmark for complex retrieval tasks—covering multi-condition filtering, multi-hop reasoning, and natural-language constraints. Contribution/Results: Our benchmark reveals severe performance degradation of state-of-the-art retrieval models under complex queries (average nDCG@10 = 0.346, R@100 = 0.587). Notably, LLM-based query rewriting—widely assumed beneficial—degrades performance even for strong retrievers, challenging prevailing assumptions. Extensive experiments across modern retrieval architectures (e.g., dense, sparse, hybrid) and LLM-augmented strategies provide reproducible evaluation protocols and critical insights for next-generation general-purpose retrieval models.
Current information retrieval systems are designed with human users in mind and struggle to accommodate search behaviors initiated by autonomous agents, leading to performance degradation and evaluation bias. To address this gap, this work proposes a systematic approach that leverages a multi-agent framework and diverse retrieval pipelines to collect agent-generated queries, retrieved documents, and reasoning traces on established benchmarks such as HotpotQA, Researchy Questions, and MS MARCO. We construct and release the first dataset specifically tailored to agentic search behavior—Agentic Search Queryset (ASQ)—alongside a supporting toolkit. This resource fills a critical void in authentic interaction data for agent-driven retrieval, enables flexible extension to new agents, retrievers, and tasks, and lays the foundation for future research in agentic information retrieval.
To address the poor semantic retrieval performance for long-tail e-commerce search queries—caused by sparse user interaction signals—this paper proposes a semantic product retrieval method tailored for low-frequency queries. The method comprises three key components: (1) an LLM-enhanced query signal augmentation mechanism to mitigate scarcity of labeled and interaction data; (2) a domain-adaptive dual-tower architecture integrating Transformer pretraining with multi-task contrastive learning—specifically optimizing query-query and query-product pair representations; and (3) model weight ensembling coupled with human-in-the-loop annotation to construct a high-quality evaluation dataset. In live A/B testing on a real-world e-commerce platform, the proposed method achieves a 3% lift in conversion rate over conventional lexical-matching recall baselines, demonstrating substantial improvements in both retrieval accuracy for long-tail queries and overall business impact.
Under information overload, the retrieval stage in recommender systems has long been underappreciated and lacks systematic investigation. This paper presents the first comprehensive survey of retrieval in industrial multi-stage recommendation pipelines, focusing on three core aspects: user-item similarity modeling, efficient indexing mechanisms (e.g., vector search and inverted indices), and training optimization techniques—including dual-tower architectures, contrastive learning, and negative sampling. We introduce a unified evaluation benchmark spanning three public datasets and integrate insights from leading industry practitioners to holistically characterize deployment practices, performance bottlenecks, and engineering challenges. Our work fills a critical gap in the systematic analysis of retrieval and provides both theoretical foundations and practical paradigms for designing accurate, efficient, and production-ready retrieval components within cascaded recommendation systems.
This study addresses the lack of standardized validation criteria for user simulators in information retrieval evaluation, which undermines the reliability of simulation outcomes. Through a systematic literature review, it proposes the first structured taxonomy of metrics specifically designed for validating simulated search queries. The work empirically analyzes the interrelationships among these metrics across four diverse datasets and, based on the findings, offers tailored validation recommendations for different application scenarios. To foster standardization and reproducibility in simulation-based evaluation, the authors also release an open-source toolkit implementing commonly used validation metrics, thereby supporting future research extension and benchmarking.
This work addresses the limitation of existing deep research agents that overlook structured web fields—such as titles, sections, and metadata—during the search-and-retrieve process, leading to redundant retrieval and excessive contextual noise. To overcome this, the authors propose SIEVE, a novel framework that introduces fielded Boolean queries (BQL) into the research agent paradigm for the first time. SIEVE implements a three-stage pipeline: search, inspect, and retrieve. It first filters candidate pages using document-level fields, then presents results as structured cards for selective inspection, and finally retrieves only the content of chosen sections. By leveraging web structure at a fine-grained level, SIEVE achieves higher accuracy than state-of-the-art baselines across three question-answering benchmarks while reducing context token consumption by 20.7%–50.6%, demonstrating both efficiency and broad applicability.
This work addresses the architectural challenges faced by industrial-scale web retrieval systems under stringent constraints of latency, scalability, and resource efficiency. It proposes a unified multi-stage abstraction termed “Retrieval-as-a-Service” (RaaS), which, for the first time, integrates infrastructure-aware components—including efficient candidate generation, embedding-based semantic matching, and resource-conscious re-ranking—into a cohesive framework. The study systematically models the impact of incorporating large language models (LLMs) on both system performance and operational overhead. By analyzing real-world production deployments, the authors uncover fundamental trade-offs between system design choices and quality-of-service (QoS) objectives, thereby offering practical, scalable, and QoS-aware architectural guidelines for building high-performance web-scale retrieval systems.
Traditional information retrieval prioritizes topical relevance, which often fails to meet the practical utility demands of downstream large language model (LLM) tasks. This work proposes a utility-centered retrieval paradigm that shifts the objective from relevance to the actual contribution of retrieved content to LLM generation quality. It introduces the first unified framework encompassing diverse utility forms—spanning LLM-agnostic and LLM-aware, as well as context-independent and context-dependent settings—and explicitly links LLM information needs with agent-based RAG mechanisms. By integrating retrieval-augmented generation, information need modeling, utility-oriented evaluation metrics, and intelligent retrieval strategies, this study redefines retrieval evaluation criteria and establishes a synergistic optimization pathway between retrieval and generation, offering both theoretical foundations and practical guidance for information retrieval in the LLM era.
Traditional information retrieval primarily emphasizes surface-level similarity between documents and queries, often overlooking their actual utility in supporting decision-making. This work proposes a novel retrieval paradigm centered on decision usefulness and introduces UsefulBench, the first benchmark dataset annotated by domain experts for both relevance and usefulness. Through systematic comparisons among classical retrieval models, large language models (LLMs), and human expert judgments, the study reveals that conventional methods strongly favor relevance, while LLMs, despite modest gains in usefulness, still fall short of replicating expert-level assessments. By establishing a new evaluation framework grounded in real-world utility, this research advances the foundation for usefulness-oriented information retrieval and provides a valuable resource for future development and assessment of retrieval systems.