LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest

📅 2025-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Manual relevance annotation in personalized search is costly and poorly scalable. Method: We propose an automated relevance assessment framework based on fine-tuned large language models (LLMs), incorporating query-document semantic matching and context-aware discrimination, trained via supervised fine-tuning on high-quality human annotations. Contribution/Results: This is the first work to achieve high inter-annotator agreement (Cohen’s κ > 0.85) between LLMs and human annotators in a large-scale production search system (Pinterest). The method triples query coverage, reduces the minimum detectable effect (MDE) in online experiments by 42%, and significantly improves statistical power and metric reliability. Our approach establishes a new industrial-grade paradigm for search evaluation—efficient, scalable, and high-fidelity.

Technology Category

Search and Optimization: Evaluation and AnalysisMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Relevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user's queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present our approach at Pinterest Search to automate relevance evaluation for online experiments using fine-tuned LLMs. We rigorously validate the alignment between LLM-generated judgments and human annotations, demonstrating that LLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging LLM-based labeling further unlocks the opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effect (MDE) in online experiment measurements.
Problem

Research questions and friction points this paper is trying to address.

Automating relevance evaluation using fine-tuned LLMs
Replacing costly human annotation with scalable AI solutions
Improving search quality metrics and experiment sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuned LLMs automate search relevance evaluation
LLM judgments align with human annotations reliably
Enables expanded query set and optimized sampling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Han Wang
Pinterest, San Francisco, CA, USA
A
Alex Whitworth
Pinterest, San Francisco, CA, USA
P
Pak Ming Cheung
Pinterest, San Francisco, CA, USA
Z
Zhenjie Zhang
Pinterest, San Francisco, CA, USA
K
Krishna Kamath
Pinterest, San Francisco, CA, USA