tune search relevance

Designs and builds models, scoring functions, and ranking systems that predict and assign relevance scores to candidate items given a query or information need, covering semantic and text-based relevance modeling and scoring. Implements evaluation pipelines and metrics, defines annotation and labeling schemes, and runs offline and online experiments to analyze, tune, and optimize relevance ranking and evaluation.

tunesearchrelevance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.

efficiency-effectiveness trade-offlong-context documentsretrieval system

Must-Read Papers

Most classic and influential ideas
View more

Analytical Search

Feb 12, 2026

Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.

analytical searchevidence fusioninformation retrieval

This work addresses the high cost and prolonged turnaround of manual relevance labeling, which hinder large-scale online search experimentation. To overcome these limitations, the study introduces vision-language models (VLMs) into industrial search relevance evaluation for the first time, establishing an automated labeling pipeline deployed in Pinterest’s online A/B experiments. The proposed approach substantially improves evaluation efficiency and coverage, enabling more granular sampling strategies and reducing the minimum detectable effect (MDE). Empirical results demonstrate strong agreement between VLM-generated relevance judgments and human annotations, confirming the method’s capacity to support high-quality, high-sensitivity assessment of search systems at scale.

human annotationpersonalized searchrelevance evaluation

This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.

accessibilitylarge language modelsmodel scale

RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback

Aug 07, 2025
AA
Abdelrahman Abdallah
🏛️ University of Innsbruck | Chungbuk National University

Current RAG and re-ranking systems lack scalable, user-centric, and multi-perspective evaluation tools. To address this, we propose the first unified platform enabling end-to-end joint evaluation of retrieval, re-ranking, and RAG. Our method innovatively integrates dual feedback mechanisms—human expert annotation and LLM-as-a-judge—supporting pairwise comparison, full-list labeling, blind voting, visualized ranking, and structured metadata collection. The platform enables fine-grained relevance annotation and question-answering quality analysis, producing high-quality, reusable, structured evaluation datasets that directly facilitate downstream tasks such as re-ranker optimization and reward modeling. All code is open-sourced, and an online demo is provided. Empirical results demonstrate significant improvements in evaluation reliability, interpretability, and engineering practicality.

Lack scalable tools for RAG and reranking evaluationNeed unified platform for multi-perspective feedback collectionRequire comparison of human and LLM-based ranking judgments

To address the high cost, low efficiency, and error-proneness of manual query-item relevance annotation in e-commerce search, this paper proposes an automated annotation framework leveraging large language models (LLMs), specifically LLaMA and GPT. We present the first systematic empirical validation that LLMs can achieve human-expert-level performance on large-scale relevance judgment tasks. Our method introduces a novel multi-strategy prompting framework integrating chain-of-thought (CoT) reasoning, in-context learning (ICL), and retrieval-augmented generation with maximal marginal relevance (RAG-MMR). Evaluated across multiple public and proprietary datasets, our approach achieves annotation accuracy within ±1.2% of human benchmarks while improving throughput by over 100×. The framework has been deployed in production for training and iterative evaluation of search ranking models, significantly reducing annotation costs and accelerating R&D cycles.

Automate query-product relevance labelingEnhance e-commerce search efficiencyReduce human-labeling cost and time

Latest Papers

What's happening recently
View more

Benchmarking Information Retrieval Models on Complex Retrieval Tasks

Sep 08, 2025
JK
Julian Killingback
🏛️ University of Massachusetts Amherst

Existing retrieval evaluation benchmarks predominantly rely on simple, single-point queries, failing to reflect model capabilities under realistic, complex retrieval scenarios involving multiple constraints and intents. Method: We introduce ComplexRetrieval-Bench—the first systematic, diverse, and realistic benchmark for complex retrieval tasks—covering multi-condition filtering, multi-hop reasoning, and natural-language constraints. Contribution/Results: Our benchmark reveals severe performance degradation of state-of-the-art retrieval models under complex queries (average nDCG@10 = 0.346, R@100 = 0.587). Notably, LLM-based query rewriting—widely assumed beneficial—degrades performance even for strong retrievers, challenging prevailing assumptions. Extensive experiments across modern retrieval architectures (e.g., dense, sparse, hybrid) and LLM-augmented strategies provide reproducible evaluation protocols and critical insights for next-generation general-purpose retrieval models.

Assessing retrieval models on diverse complex tasks with realistic settingsEvaluating performance on queries containing multiple constraints and requirementsMeasuring impact of LLM-based query expansion on retrieval quality

Manual relevance annotation in personalized search is costly and poorly scalable. Method: We propose an automated relevance assessment framework based on fine-tuned large language models (LLMs), incorporating query-document semantic matching and context-aware discrimination, trained via supervised fine-tuning on high-quality human annotations. Contribution/Results: This is the first work to achieve high inter-annotator agreement (Cohen’s κ > 0.85) between LLMs and human annotators in a large-scale production search system (Pinterest). The method triples query coverage, reduces the minimum detectable effect (MDE) in online experiments by 42%, and significantly improves statistical power and metric reliability. Our approach establishes a new industrial-grade paradigm for search evaluation—efficient, scalable, and high-fidelity.

Automating relevance evaluation using fine-tuned LLMsImproving search quality metrics and experiment sensitivityReplacing costly human annotation with scalable AI solutions

This work addresses the semantic gap between user queries and product descriptions in e-commerce search by proposing a multi-task, multi-stage query rewriting framework based on large language models. The approach uniquely integrates explicit relevance modeling into the rewriting process, combining supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO)—a reinforcement learning algorithm tailored to business objectives—to jointly optimize query rewriting, relevance estimation, and user conversion. Experiments leveraging JD.com’s pretrained large language model demonstrate significant improvements in both offline relevance metrics and online user conversion rate (UCVR) in A/B tests. The method has been deployed on JD.com’s search platform since August 2025.

e-commerce searchlexical gapquery rewriting

This work proposes a calibrated model cascade framework to enable low-cost, efficient generation of large-scale, high-quality search relevance annotations. The approach routes queries through a sequence of progressively specialized fine-tuned classifiers, decomposing the annotation task to improve fine-tuning accuracy by 20 percentage points. The cascade architecture reduces computational overhead by approximately 50% with negligible loss in precision. Furthermore, the framework incorporates a class-wise monotonic calibration strategy that yields a statistically significant accuracy gain of +0.6 points. Validated across six production scenarios and applied to over 150 million annotations, the system substantially accelerates offline experimentation cycles while maintaining high annotation quality.

cost-efficiencyhuman labelinglarge-scale evaluation

Hot Scholars

XC

Xueqi Cheng

Ph.D. student, Florida State University
Data miningLLMGNNComputational social science
EC

Enhong Chen

University of Science and Technology of China
data miningrecommender systemmachine learning
JL

Jimmy Lin

University of Waterloo
information retrievalnatural language processingdata managementbig data
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
ZL

Zhenghao Liu

Northeastern University
NLPInformation Retrieval