Score
Designs, builds, or analyzes systems and components that locate and retrieve relevant items in response to user queries — e.g., indexes, query parsers, retrieval algorithms, ranking and scoring models, and result presentation. Work also covers relevance estimation and evaluation, search pipeline infrastructure (latency, scalability), and interaction features such as query suggestions and filtering.
Prior surveys on information retrieval (IR) models conflate architectural design with training methodologies, obscuring the intrinsic evolution of structural innovations in relevance modeling. Method: We systematically trace the architectural progression of IR models—spanning backbone feature extractors and end-to-end relevance modeling—from classical BM25 through CNN/RNN-based rankers to modern BERT dual-encoder and interaction-based architectures, ColBERT, Cross-Encoders, and LLM-based retrievers—explicitly decoupling architecture from training strategy. Contribution/Results: We propose the first longitudinal IR-specific architectural taxonomy, explicitly addressing scalability and adaptability challenges in multimodal, multilingual, and emerging application scenarios. Our framework provides an actionable technology roadmap for industrial system selection and rigorously identifies open research questions and future directions for the academic community.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This work addresses the architectural challenges faced by industrial-scale web retrieval systems under stringent constraints of latency, scalability, and resource efficiency. It proposes a unified multi-stage abstraction termed “Retrieval-as-a-Service” (RaaS), which, for the first time, integrates infrastructure-aware components—including efficient candidate generation, embedding-based semantic matching, and resource-conscious re-ranking—into a cohesive framework. The study systematically models the impact of incorporating large language models (LLMs) on both system performance and operational overhead. By analyzing real-world production deployments, the authors uncover fundamental trade-offs between system design choices and quality-of-service (QoS) objectives, thereby offering practical, scalable, and QoS-aware architectural guidelines for building high-performance web-scale retrieval systems.
Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.
Current information retrieval systems are designed with human users in mind and struggle to accommodate search behaviors initiated by autonomous agents, leading to performance degradation and evaluation bias. To address this gap, this work proposes a systematic approach that leverages a multi-agent framework and diverse retrieval pipelines to collect agent-generated queries, retrieved documents, and reasoning traces on established benchmarks such as HotpotQA, Researchy Questions, and MS MARCO. We construct and release the first dataset specifically tailored to agentic search behavior—Agentic Search Queryset (ASQ)—alongside a supporting toolkit. This resource fills a critical void in authentic interaction data for agent-driven retrieval, enables flexible extension to new agents, retrievers, and tasks, and lays the foundation for future research in agentic information retrieval.
To address scalability bottlenecks in Text-to-SQL for enterprise-scale databases, this paper proposes a domain-agnostic, retrieval-augmented schema linking framework. Methodologically, it introduces a modular semantic indexing architecture that decouples database schemas and metadata into fine-grained semantic units for independent vectorization; further, it designs a table-level prioritized identification mechanism coupled with column-level contextual fusion, integrating chunked indexing, semantic retrieval, recall re-ranking, and context budget control. Crucially, the approach fully leverages intrinsic semantic cues from raw metadata without requiring domain-specific fine-tuning, enabling high-precision schema matching. Experimental results demonstrate that our method consistently outperforms mainstream baselines across heterogeneous, multi-source data catalogs—achieving both high recall and high accuracy. Its plug-and-play compatibility facilitates seamless enterprise deployment.
This work addresses the limitation of existing deep research agents that overlook structured web fields—such as titles, sections, and metadata—during the search-and-retrieve process, leading to redundant retrieval and excessive contextual noise. To overcome this, the authors propose SIEVE, a novel framework that introduces fielded Boolean queries (BQL) into the research agent paradigm for the first time. SIEVE implements a three-stage pipeline: search, inspect, and retrieve. It first filters candidate pages using document-level fields, then presents results as structured cards for selective inspection, and finally retrieves only the content of chosen sections. By leveraging web structure at a fine-grained level, SIEVE achieves higher accuracy than state-of-the-art baselines across three question-answering benchmarks while reducing context token consumption by 20.7%–50.6%, demonstrating both efficiency and broad applicability.
This work proposes the concept of “social relevance” to address the limitations of existing search relevance models, which primarily focus on topical matching and often fail to identify or mitigate harmful content such as misinformation and discriminatory material, thereby neglecting broader societal interests. Through conceptual analysis and a three-dimensional modeling framework encompassing system, user, and societal perspectives, the study clarifies the definition, boundaries, and distinction of social relevance from traditional information quality metrics. By transcending conventional relevance paradigms, this framework provides a theoretical foundation for designing retrieval systems that integrate ethical values and social welfare, ultimately advancing search engines toward a value-driven paradigm.
This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.
This work addresses the semantic gap between user queries and product descriptions in e-commerce search by proposing a multi-task, multi-stage query rewriting framework based on large language models. The approach uniquely integrates explicit relevance modeling into the rewriting process, combining supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO)—a reinforcement learning algorithm tailored to business objectives—to jointly optimize query rewriting, relevance estimation, and user conversion. Experiments leveraging JD.com’s pretrained large language model demonstrate significant improvements in both offline relevance metrics and online user conversion rate (UCVR) in A/B tests. The method has been deployed on JD.com’s search platform since August 2025.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
Traditional information retrieval primarily emphasizes surface-level similarity between documents and queries, often overlooking their actual utility in supporting decision-making. This work proposes a novel retrieval paradigm centered on decision usefulness and introduces UsefulBench, the first benchmark dataset annotated by domain experts for both relevance and usefulness. Through systematic comparisons among classical retrieval models, large language models (LLMs), and human expert judgments, the study reveals that conventional methods strongly favor relevance, while LLMs, despite modest gains in usefulness, still fall short of replicating expert-level assessments. By establishing a new evaluation framework grounded in real-world utility, this research advances the foundation for usefulness-oriented information retrieval and provides a valuable resource for future development and assessment of retrieval systems.