Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

📅 2025-07-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
For subjective, multi-dimensional annotation tasks—such as search query clarification—current large language models (LLMs) still fall short of human-level performance in automated labeling, necessitating robust human-in-the-loop mechanisms. Method: We propose a lightweight human-in-the-loop annotation framework that leverages multi-LLM ensemble reasoning and confidence calibration to dynamically identify low-confidence and inter-model disagreement samples, thereby triggering targeted human review. This enables the construction of high-quality, multi-dimensional labeled datasets through a quality-controllable hybrid annotation pipeline. Contribution/Results: Experiments demonstrate that our approach maintains annotation consistency and reliability while reducing human effort by up to 45%. It significantly improves annotation efficiency and scalability, offering a cost-effective, robust paradigm for deploying LLMs in complex evaluation scenarios.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Humans and AI: Human-in-the-loop Machine LearningNatural Language Processing: (Large) Language Models

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Despite growing interest in using large language models (LLMs) to automate annotation, their effectiveness in complex, nuanced, and multi-dimensional labelling tasks remains relatively underexplored. This study focuses on annotation for the search clarification task, leveraging a high-quality, multi-dimensional dataset that includes five distinct fine-grained annotation subtasks. Although LLMs have shown impressive capabilities in general settings, our study reveals that even state-of-the-art models struggle to replicate human-level performance in subjective or fine-grained evaluation tasks. Through a systematic assessment, we demonstrate that LLM predictions are often inconsistent, poorly calibrated, and highly sensitive to prompt variations. To address these limitations, we propose a simple yet effective human-in-the-loop (HITL) workflow that uses confidence thresholds and inter-model disagreement to selectively involve human review. Our findings show that this lightweight intervention significantly improves annotation reliability while reducing human effort by up to 45%, offering a relatively scalable and cost-effective yet accurate path forward for deploying LLMs in real-world evaluation settings.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs' effectiveness in complex multi-dimensional annotation tasks
Addressing LLMs' inconsistency and sensitivity in fine-grained labeling
Proposing human-in-the-loop workflow to improve annotation reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-in-the-loop workflow for reliable annotations
Confidence thresholds to reduce human effort
Inter-model disagreement for selective human review