🤖 AI Summary
For subjective, multi-dimensional annotation tasks—such as search query clarification—current large language models (LLMs) still fall short of human-level performance in automated labeling, necessitating robust human-in-the-loop mechanisms.
Method: We propose a lightweight human-in-the-loop annotation framework that leverages multi-LLM ensemble reasoning and confidence calibration to dynamically identify low-confidence and inter-model disagreement samples, thereby triggering targeted human review. This enables the construction of high-quality, multi-dimensional labeled datasets through a quality-controllable hybrid annotation pipeline.
Contribution/Results: Experiments demonstrate that our approach maintains annotation consistency and reliability while reducing human effort by up to 45%. It significantly improves annotation efficiency and scalability, offering a cost-effective, robust paradigm for deploying LLMs in complex evaluation scenarios.
📝 Abstract
Despite growing interest in using large language models (LLMs) to automate annotation, their effectiveness in complex, nuanced, and multi-dimensional labelling tasks remains relatively underexplored. This study focuses on annotation for the search clarification task, leveraging a high-quality, multi-dimensional dataset that includes five distinct fine-grained annotation subtasks. Although LLMs have shown impressive capabilities in general settings, our study reveals that even state-of-the-art models struggle to replicate human-level performance in subjective or fine-grained evaluation tasks. Through a systematic assessment, we demonstrate that LLM predictions are often inconsistent, poorly calibrated, and highly sensitive to prompt variations. To address these limitations, we propose a simple yet effective human-in-the-loop (HITL) workflow that uses confidence thresholds and inter-model disagreement to selectively involve human review. Our findings show that this lightweight intervention significantly improves annotation reliability while reducing human effort by up to 45%, offering a relatively scalable and cost-effective yet accurate path forward for deploying LLMs in real-world evaluation settings.