Score
Designs and builds models, scoring functions, and ranking systems that predict and assign relevance scores to candidate items given a query or information need, covering semantic and text-based relevance modeling and scoring. Implements evaluation pipelines and metrics, defines annotation and labeling schemes, and runs offline and online experiments to analyze, tune, and optimize relevance ranking and evaluation.
Prior surveys on information retrieval (IR) models conflate architectural design with training methodologies, obscuring the intrinsic evolution of structural innovations in relevance modeling. Method: We systematically trace the architectural progression of IR models—spanning backbone feature extractors and end-to-end relevance modeling—from classical BM25 through CNN/RNN-based rankers to modern BERT dual-encoder and interaction-based architectures, ColBERT, Cross-Encoders, and LLM-based retrievers—explicitly decoupling architecture from training strategy. Contribution/Results: We propose the first longitudinal IR-specific architectural taxonomy, explicitly addressing scalability and adaptability challenges in multimodal, multilingual, and emerging application scenarios. Our framework provides an actionable technology roadmap for industrial system selection and rigorously identifies open research questions and future directions for the academic community.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.
This work addresses the high cost and prolonged turnaround of manual relevance labeling, which hinder large-scale online search experimentation. To overcome these limitations, the study introduces vision-language models (VLMs) into industrial search relevance evaluation for the first time, establishing an automated labeling pipeline deployed in Pinterest’s online A/B experiments. The proposed approach substantially improves evaluation efficiency and coverage, enabling more granular sampling strategies and reducing the minimum detectable effect (MDE). Empirical results demonstrate strong agreement between VLM-generated relevance judgments and human annotations, confirming the method’s capacity to support high-quality, high-sensitivity assessment of search systems at scale.
This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.
Current RAG and re-ranking systems lack scalable, user-centric, and multi-perspective evaluation tools. To address this, we propose the first unified platform enabling end-to-end joint evaluation of retrieval, re-ranking, and RAG. Our method innovatively integrates dual feedback mechanisms—human expert annotation and LLM-as-a-judge—supporting pairwise comparison, full-list labeling, blind voting, visualized ranking, and structured metadata collection. The platform enables fine-grained relevance annotation and question-answering quality analysis, producing high-quality, reusable, structured evaluation datasets that directly facilitate downstream tasks such as re-ranker optimization and reward modeling. All code is open-sourced, and an online demo is provided. Empirical results demonstrate significant improvements in evaluation reliability, interpretability, and engineering practicality.
To address the high cost, low efficiency, and error-proneness of manual query-item relevance annotation in e-commerce search, this paper proposes an automated annotation framework leveraging large language models (LLMs), specifically LLaMA and GPT. We present the first systematic empirical validation that LLMs can achieve human-expert-level performance on large-scale relevance judgment tasks. Our method introduces a novel multi-strategy prompting framework integrating chain-of-thought (CoT) reasoning, in-context learning (ICL), and retrieval-augmented generation with maximal marginal relevance (RAG-MMR). Evaluated across multiple public and proprietary datasets, our approach achieves annotation accuracy within ±1.2% of human benchmarks while improving throughput by over 100×. The framework has been deployed in production for training and iterative evaluation of search ranking models, significantly reducing annotation costs and accelerating R&D cycles.
Existing retrieval evaluation benchmarks predominantly rely on simple, single-point queries, failing to reflect model capabilities under realistic, complex retrieval scenarios involving multiple constraints and intents. Method: We introduce ComplexRetrieval-Bench—the first systematic, diverse, and realistic benchmark for complex retrieval tasks—covering multi-condition filtering, multi-hop reasoning, and natural-language constraints. Contribution/Results: Our benchmark reveals severe performance degradation of state-of-the-art retrieval models under complex queries (average nDCG@10 = 0.346, R@100 = 0.587). Notably, LLM-based query rewriting—widely assumed beneficial—degrades performance even for strong retrievers, challenging prevailing assumptions. Extensive experiments across modern retrieval architectures (e.g., dense, sparse, hybrid) and LLM-augmented strategies provide reproducible evaluation protocols and critical insights for next-generation general-purpose retrieval models.
Manual relevance annotation in personalized search is costly and poorly scalable. Method: We propose an automated relevance assessment framework based on fine-tuned large language models (LLMs), incorporating query-document semantic matching and context-aware discrimination, trained via supervised fine-tuning on high-quality human annotations. Contribution/Results: This is the first work to achieve high inter-annotator agreement (Cohen’s κ > 0.85) between LLMs and human annotators in a large-scale production search system (Pinterest). The method triples query coverage, reduces the minimum detectable effect (MDE) in online experiments by 42%, and significantly improves statistical power and metric reliability. Our approach establishes a new industrial-grade paradigm for search evaluation—efficient, scalable, and high-fidelity.
This work addresses the semantic gap between user queries and product descriptions in e-commerce search by proposing a multi-task, multi-stage query rewriting framework based on large language models. The approach uniquely integrates explicit relevance modeling into the rewriting process, combining supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO)—a reinforcement learning algorithm tailored to business objectives—to jointly optimize query rewriting, relevance estimation, and user conversion. Experiments leveraging JD.com’s pretrained large language model demonstrate significant improvements in both offline relevance metrics and online user conversion rate (UCVR) in A/B tests. The method has been deployed on JD.com’s search platform since August 2025.
This work proposes a calibrated model cascade framework to enable low-cost, efficient generation of large-scale, high-quality search relevance annotations. The approach routes queries through a sequence of progressively specialized fine-tuned classifiers, decomposing the annotation task to improve fine-tuning accuracy by 20 percentage points. The cascade architecture reduces computational overhead by approximately 50% with negligible loss in precision. Furthermore, the framework incorporates a class-wise monotonic calibration strategy that yields a statistically significant accuracy gain of +0.6 points. Validated across six production scenarios and applied to over 150 million annotations, the system substantially accelerates offline experimentation cycles while maintaining high annotation quality.