retrieval

Designs, implements, and evaluates systems and pipelines that locate and return relevant items from a dataset or corpus in response to a query or signal. This includes building indexes and retrieval models (e.g., inverted indexes, vector search), ranking and relevance-scoring methods, and measuring performance with metrics such as recall and precision.

retrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
3.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Analytical Search

Feb 12, 2026

Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.

analytical searchevidence fusioninformation retrieval

Traditional information retrieval primarily emphasizes surface-level similarity between documents and queries, often overlooking their actual utility in supporting decision-making. This work proposes a novel retrieval paradigm centered on decision usefulness and introduces UsefulBench, the first benchmark dataset annotated by domain experts for both relevance and usefulness. Through systematic comparisons among classical retrieval models, large language models (LLMs), and human expert judgments, the study reveals that conventional methods strongly favor relevance, while LLMs, despite modest gains in usefulness, still fall short of replicating expert-level assessments. By establishing a new evaluation framework grounded in real-world utility, this research advances the foundation for usefulness-oriented information retrieval and provides a valuable resource for future development and assessment of retrieval systems.

decision-useful informationdomain-specific expertiseinformation retrieval

Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models

Oct 31, 2024
JZ
Jianqun Zhou
🏛️ Eastern Institute of Technology | Huazhong University of Science and Technology | Salesforce Research

Existing retrieval models exhibit limited capability in adhering to user-specified document-level instructions—such as target audience, output format, or language preferences. Method: We propose InfoSearch, the first instruction-following document-level retrieval benchmark, introducing two novel evaluation metrics: Strict Instruction Compliance Rate (SICR) and Weighted Instruction Sensitivity Evaluation (WISE). We further design an LLM-driven, instruction-aware retrieval framework that integrates dense retrieval with attribute-aware re-ranking to support multi-dimensional constraint modeling. Contribution/Results: Empirical evaluation reveals consistently low instruction compliance across mainstream retrieval models. While fine-tuning and scaling improve performance, substantial gaps remain relative to practical deployment requirements. This work establishes a systematic evaluation framework and technical foundation for instruction-aware retrieval.

Develops InfoSearch benchmark for six document-level attributes.Evaluates instruction-following in retrieval models beyond content relevance.Introduces SICR and WISE metrics to assess instruction responsiveness.

This work addresses the limitation of existing deep research agents that overlook structured web fields—such as titles, sections, and metadata—during the search-and-retrieve process, leading to redundant retrieval and excessive contextual noise. To overcome this, the authors propose SIEVE, a novel framework that introduces fielded Boolean queries (BQL) into the research agent paradigm for the first time. SIEVE implements a three-stage pipeline: search, inspect, and retrieve. It first filters candidate pages using document-level fields, then presents results as structured cards for selective inspection, and finally retrieves only the content of chosen sections. By leveraging web structure at a fine-grained level, SIEVE achieves higher accuracy than state-of-the-art baselines across three question-answering benchmarks while reducing context token consumption by 20.7%–50.6%, demonstrating both efficiency and broad applicability.

Boolean retrievaldeep-research agentsdocument structure

This study addresses the challenge of supporting long-term interactive exploratory search, where user intent is often ambiguous and dynamically evolving, making it difficult for existing retrieval models to effectively respond to fine-grained instructions. The work presents the first systematic evaluation of instruction-tuned large language models (LLMs) in aspect-oriented, seed-guided exploratory search tasks, assessing both instruction-following fidelity and ranking relevance using an expert-annotated test set. Experiments employ both fine-tuned instruction-tuned LLMs and general-purpose LLMs integrated with Pairwise Ranking Prompting for document ranking. Results show that while the best-performing model achieves superior ranking relevance compared to non-instruction-aware baselines, its ability to follow instructions does not improve correspondingly—exhibiting insensitivity or even counterintuitive behaviors in response to user directives. This reveals a critical limitation of current instruction-based retrieval approaches in interactive exploratory settings.

aspect-conditional explorationexploratory searchinstructed retrieval

Latest Papers

What's happening recently
View more

This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.

accessibilitylarge language modelsmodel scale

This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.

cascading pipelinesexperiment efficiencyinformation retrieval

This work proposes the concept of “social relevance” to address the limitations of existing search relevance models, which primarily focus on topical matching and often fail to identify or mitigate harmful content such as misinformation and discriminatory material, thereby neglecting broader societal interests. Through conceptual analysis and a three-dimensional modeling framework encompassing system, user, and societal perspectives, the study clarifies the definition, boundaries, and distinction of social relevance from traditional information quality metrics. By transcending conventional relevance paradigms, this framework provides a theoretical foundation for designing retrieval systems that integrate ethical values and social welfare, ultimately advancing search engines toward a value-driven paradigm.

ethical outcomesharmful contentinformation retrieval

Traditional keyword-based code retrieval struggles to meet the demands of natural language queries, intent understanding, and code quality assessment. This work proposes a hybrid retrieval system that integrates semantic search with large language model (LLM)-generated quality metadata, supporting four query modes: semantic, quality-filtered, hybrid, and automatic routing. The approach innovatively incorporates function-level code slicing, text-code embeddings, and ChromaDB vector storage, and—novelly—leverages LLM-generated quality scores for dynamic query routing. Experiments on a C-language educational code corpus demonstrate strong performance: semantic retrieval achieves nDCG@5 of 0.820 and Success@5 of 0.800; automatic routing attains 100% accuracy; and in 9 out of 12 cases, LLM-predicted quality scores deviate by no more than one point from human evaluations.

code qualityimplementation intentnatural language queries

This work addresses the architectural challenges faced by industrial-scale web retrieval systems under stringent constraints of latency, scalability, and resource efficiency. It proposes a unified multi-stage abstraction termed “Retrieval-as-a-Service” (RaaS), which, for the first time, integrates infrastructure-aware components—including efficient candidate generation, embedding-based semantic matching, and resource-conscious re-ranking—into a cohesive framework. The study systematically models the impact of incorporating large language models (LLMs) on both system performance and operational overhead. By analyzing real-world production deployments, the authors uncover fundamental trade-offs between system design choices and quality-of-service (QoS) objectives, thereby offering practical, scalable, and QoS-aware architectural guidelines for building high-performance web-scale retrieval systems.

industrial retrieval pipelineslatency requirementsRetrieval-as-a-Service