search engines

Designs, implements, and evaluates systems that collect, index, and retrieve information in response to user queries, including components such as crawlers, document indexes, query parsers, ranking algorithms, and retrieval pipelines. Builds and tunes relevance signals, query understanding, result presentation, and system-level concerns (latency, scalability, and evaluation) to enable effective information access.

searchengines

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current information retrieval systems are designed with human users in mind and struggle to accommodate search behaviors initiated by autonomous agents, leading to performance degradation and evaluation bias. To address this gap, this work proposes a systematic approach that leverages a multi-agent framework and diverse retrieval pipelines to collect agent-generated queries, retrieved documents, and reasoning traces on established benchmarks such as HotpotQA, Researchy Questions, and MS MARCO. We construct and release the first dataset specifically tailored to agentic search behavior—Agentic Search Queryset (ASQ)—alongside a supporting toolkit. This resource fills a critical void in authentic interaction data for agent-driven retrieval, enables flexible extension to new agents, retrievers, and tasks, and lays the foundation for future research in agentic information retrieval.

agent behavioragentic searchdataset gap

This work addresses the architectural challenges faced by industrial-scale web retrieval systems under stringent constraints of latency, scalability, and resource efficiency. It proposes a unified multi-stage abstraction termed “Retrieval-as-a-Service” (RaaS), which, for the first time, integrates infrastructure-aware components—including efficient candidate generation, embedding-based semantic matching, and resource-conscious re-ranking—into a cohesive framework. The study systematically models the impact of incorporating large language models (LLMs) on both system performance and operational overhead. By analyzing real-world production deployments, the authors uncover fundamental trade-offs between system design choices and quality-of-service (QoS) objectives, thereby offering practical, scalable, and QoS-aware architectural guidelines for building high-performance web-scale retrieval systems.

industrial retrieval pipelineslatency requirementsRetrieval-as-a-Service

This work addresses the limitation of existing deep research agents that overlook structured web fields—such as titles, sections, and metadata—during the search-and-retrieve process, leading to redundant retrieval and excessive contextual noise. To overcome this, the authors propose SIEVE, a novel framework that introduces fielded Boolean queries (BQL) into the research agent paradigm for the first time. SIEVE implements a three-stage pipeline: search, inspect, and retrieve. It first filters candidate pages using document-level fields, then presents results as structured cards for selective inspection, and finally retrieves only the content of chosen sections. By leveraging web structure at a fine-grained level, SIEVE achieves higher accuracy than state-of-the-art baselines across three question-answering benchmarks while reducing context token consumption by 20.7%–50.6%, demonstrating both efficiency and broad applicability.

Boolean retrievaldeep-research agentsdocument structure

Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models

Oct 31, 2024
JZ
Jianqun Zhou
🏛️ Eastern Institute of Technology | Huazhong University of Science and Technology | Salesforce Research

Existing retrieval models exhibit limited capability in adhering to user-specified document-level instructions—such as target audience, output format, or language preferences. Method: We propose InfoSearch, the first instruction-following document-level retrieval benchmark, introducing two novel evaluation metrics: Strict Instruction Compliance Rate (SICR) and Weighted Instruction Sensitivity Evaluation (WISE). We further design an LLM-driven, instruction-aware retrieval framework that integrates dense retrieval with attribute-aware re-ranking to support multi-dimensional constraint modeling. Contribution/Results: Empirical evaluation reveals consistently low instruction compliance across mainstream retrieval models. While fine-tuning and scaling improve performance, substantial gaps remain relative to practical deployment requirements. This work establishes a systematic evaluation framework and technical foundation for instruction-aware retrieval.

Develops InfoSearch benchmark for six document-level attributes.Evaluates instruction-following in retrieval models beyond content relevance.Introduces SICR and WISE metrics to assess instruction responsiveness.

This work addresses a critical limitation of current large language model–based retrieval agents, which heavily rely on search engine indexes and consequently fail to access dynamic web pages, embedded files, and other unindexed content, resulting in significant blind spots. To tackle this challenge, the study formally defines the Unindexed Information Seeking (UIS) problem and introduces UIS-Digger, a multi-agent framework that employs a dual-mode browsing mechanism to simultaneously explore web pages and parse embedded documents, actively uncovering previously inaccessible information sources. The authors construct UIS-QA, the first benchmark dedicated to UIS, and optimize their system using a 30B-parameter language model enhanced through both supervised and reinforcement fine-tuning. Experimental results demonstrate that UIS-Digger achieves 27.27% accuracy on UIS-QA, substantially outperforming stronger baselines—including ensembles incorporating GPT-4.1—thereby validating the efficacy of proactive interaction with unindexed sources.

Dynamic web contentInformation-seeking agentsNon-indexed data

Latest Papers

What's happening recently
View more

Analytical Search

Feb 12, 2026

Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.

analytical searchevidence fusioninformation retrieval

This work proposes the concept of “social relevance” to address the limitations of existing search relevance models, which primarily focus on topical matching and often fail to identify or mitigate harmful content such as misinformation and discriminatory material, thereby neglecting broader societal interests. Through conceptual analysis and a three-dimensional modeling framework encompassing system, user, and societal perspectives, the study clarifies the definition, boundaries, and distinction of social relevance from traditional information quality metrics. By transcending conventional relevance paradigms, this framework provides a theoretical foundation for designing retrieval systems that integrate ethical values and social welfare, ultimately advancing search engines toward a value-driven paradigm.

ethical outcomesharmful contentinformation retrieval

This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.

component-wise optimizationdense retrievalmulti-stage training

Traditional keyword-based code retrieval struggles to meet the demands of natural language queries, intent understanding, and code quality assessment. This work proposes a hybrid retrieval system that integrates semantic search with large language model (LLM)-generated quality metadata, supporting four query modes: semantic, quality-filtered, hybrid, and automatic routing. The approach innovatively incorporates function-level code slicing, text-code embeddings, and ChromaDB vector storage, and—novelly—leverages LLM-generated quality scores for dynamic query routing. Experiments on a C-language educational code corpus demonstrate strong performance: semantic retrieval achieves nDCG@5 of 0.820 and Success@5 of 0.800; automatic routing attains 100% accuracy; and in 9 out of 12 cases, LLM-predicted quality scores deviate by no more than one point from human evaluations.

code qualityimplementation intentnatural language queries

Existing retrieval evaluation metrics, such as precision and recall, struggle to assess the breadth of information coverage in retrieved results, particularly in retrieval-augmented generation (RAG) scenarios where it is critical to capture diverse key information. To address this limitation, this work introduces CoverageBench, the first cross-task, multi-domain benchmark specifically designed for evaluating information coverage. CoverageBench integrates topics, fine-grained information nuggets, relevance labels, and baseline rankings, moving beyond traditional document-level relevance paradigms. Released via Hugging Face Datasets, the benchmark enables reproducible, quantitative evaluation of the diversity and comprehensiveness of retrieved information, establishing a standardized platform for advancing research on information coverage in retrieval systems.

ad hoc retrievalinformation coverageRAG