human-in-the-loop annotation

Design, build, and evaluate annotation systems and workflows that integrate human annotators and LLMs to produce structured labels, metadata, and human-editable prelabels; this includes schema-constrained and prompt-driven preannotation, calibrated and zero-shot LLM labeling, bootstrapping pipelines, and multi-stage diagnostic traces or candidate explanations for human review. Also develop and analyze crowdsourcing and annotator-management processes, validation and quality-control protocols, interfaces for hybrid human–LLM collaboration, and empirical studies and metrics that quantify accuracy, agreement, speed-up, and per-item labeling effort.

human-in-the-loopannotation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

May 02, 2024
MP
Maja Pavlovic
🏛️ Queen Mary University of London | University of Utrecht

This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.

Addressing limitations like bias and prompt sensitivityComparing human and GPT-generated opinion distributionsEvaluating LLMs' effectiveness in data annotation tasks

Must-Read Papers

Most classic and influential ideas
View more

From Human Annotation to LLMs: SILICON Annotation Workflow for Management Research

Dec 19, 2024
XC
Xiang Cheng
🏛️ University of Maryland | New York University

In management research, unstructured text annotation has long relied on crowdsourced human labor, while large language models (LLMs) offer efficiency and cost advantages but lack a systematic, reproducible framework for evaluating their applicability. Method: We propose SILICON, an LLM-powered text annotation workflow tailored to management research, integrating structured annotation guideline design, expert-derived baseline construction, iterative prompt optimization, and multi-model cross-validation—introducing, for the first time, a regression-based method for comparing LLM outputs. Contribution/Results: Validated via Krippendorff’s α reliability analysis and seven empirical case studies, SILICON demonstrates high agreement between LLM and expert annotations in single-label tasks (α > 0.8), but markedly reduced consistency in multi-label classification. Results confirm that expert baselines outperform crowdsourced annotations and that multi-model evaluation is indispensable. We publicly release a comprehensive practice guide and end-to-end implementation code, addressing a critical methodological gap in LLM-assisted qualitative research.

Reproducible processes for LLM annotation integrationSystematic workflow for LLM-based text annotationValidation of LLMs in management research tasks

Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

Jul 01, 2025
LT
Leila Tavakoli
🏛️ Services Australia | University of Massachusetts Amherst

For subjective, multi-dimensional annotation tasks—such as search query clarification—current large language models (LLMs) still fall short of human-level performance in automated labeling, necessitating robust human-in-the-loop mechanisms. Method: We propose a lightweight human-in-the-loop annotation framework that leverages multi-LLM ensemble reasoning and confidence calibration to dynamically identify low-confidence and inter-model disagreement samples, thereby triggering targeted human review. This enables the construction of high-quality, multi-dimensional labeled datasets through a quality-controllable hybrid annotation pipeline. Contribution/Results: Experiments demonstrate that our approach maintains annotation consistency and reliability while reducing human effort by up to 45%. It significantly improves annotation efficiency and scalability, offering a cost-effective, robust paradigm for deploying LLMs in complex evaluation scenarios.

Addressing LLMs' inconsistency and sensitivity in fine-grained labelingEvaluating LLMs' effectiveness in complex multi-dimensional annotation tasksProposing human-in-the-loop workflow to improve annotation reliability

Can LLMs Replace Manual Annotation of Software Engineering Artifacts?

Aug 10, 2024
TA
Toufique Ahmed
🏛️ University of California, Davis | Singapore Management University | University of Stuttgart

This study investigates whether large language models (LLMs) can reliably replace human annotators to mitigate the high cost and logistical complexity of human-subject studies in software engineering innovation evaluation. We systematically evaluate six state-of-the-art LLMs across ten code-related annotation tasks—including code summary quality assessment and defect repair judgment—using five public datasets. Methodologically, we propose *inter-model agreement* as a novel task-adaptability predictor and integrate confidence-threshold filtering to identify samples safe for LLM-only annotation, thereby establishing a hybrid human–LLM evaluation paradigm. Results show that LLMs achieve or approach human inter-annotator agreement (Krippendorff’s α ≥ 0.8) on multiple tasks; inter-model agreement strongly predicts task feasibility (AUC = 0.92); and confidence-based filtering raises replacement accuracy to 94.3%.

LLMs replace human annotationModel-model agreement predictorSoftware engineering artifact evaluation

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Oct 24, 2024
ON
Omer Nahum
🏛️ Technion - Institute of Technology | Google Research

Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.

Comparing annotation quality from experts, crowdsourcing, and LLMsDetecting label errors in NLP benchmark datasetsMitigating mislabeled data effects on model performance

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Jun 26, 2024
AB
Anna Bavaresco
🏛️ University of Amsterdam | University of Trento | University of Copenhagen | Utrecht University | ETH Zürich | Saarland University | Universidade de Lisboa | LMU Munich | University of Potsdam | Heriot-Watt University | Unbabel | MCML

This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.

Assessing validity of LLMs replacing human judges in NLP evaluationsEvaluating reproducibility of proprietary vs open-weight LLM modelsMeasuring variance in LLM performance across diverse NLP tasks

Latest Papers

What's happening recently
View more

This study addresses the limitations of large language models (LLMs) in annotating complex social science constructs—such as climate mitigation pessimism—where autonomous labeling often yields suboptimal quality. To overcome this, the authors propose AnnotateThis, a human-centered interactive annotation system that introduces an innovative “LLM grounding” paradigm, deeply integrating expert knowledge into the LLM annotation pipeline. The system enables iterative co-evolution of conceptual definitions and model refinement through human–AI collaboration, interactive visualizations, and dynamic prompt optimization, functioning effectively both with and without ground-truth labels. Empirical evaluation demonstrates that, in labeled settings, AnnotateThis achieves a 0.15 improvement in F-Measure and a 0.23 gain in accuracy, significantly outperforming existing fully automated approaches.

climate change mitigation pessimismcomputational social sciencedata annotation

This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.

annotation reportingannotation validityhuman annotation

This work addresses systematic biases, low agreement with human experts, irreproducibility, and data privacy risks inherent in large language models for automated annotation by proposing a highly aligned, deterministic, and open-source labeling framework based on a 4-bit quantized 1.7B-parameter small language model. Through task-aligned fine-tuning, a multidimensional scoring mechanism, and strategies involving data augmentation and regularization, the framework achieves substantially improved annotation quality with only limited human-annotated data. Experimental results demonstrate that the method outperforms the current best closed-source large models by 0.23 in Krippendorff’s α agreement metric and exhibits strong generalization across tasks such as sentiment classification. Notably, this study provides the first evidence that a quantized small model, when properly aligned via fine-tuning, can surpass state-of-the-art closed-source models in annotation consistency.

annotation alignmentdata privacyhuman-LLM disagreement

Current AI evaluation practices relying on human judgment are susceptible to anchoring effects and lack scalability, limiting their ability to provide reliable quality signals. This work proposes a human–AI collaborative evaluation framework in which humans focus on identifying salient information units—referred to as “nuggets”—and making value judgments, while large language models (LLMs) efficiently match model outputs against these nuggets. The approach integrates human oversight and automated scoring through an interactive annotation tool, a three-stage workflow, and an exportable nugget repository. By structuring human input around discrete, reusable semantic units, the framework substantially improves evaluation consistency, scalability, and accountability, thereby enhancing the reliability of LLM-as-a-Judge paradigms.

accountable evaluationhuman judgmentHuman-in-the-Loop

This work addresses the prevailing limitation in large language model (LLM) development, wherein human values are typically incorporated only post-training, lacking systematic integration across the model’s entire lifecycle. To bridge this gap, the paper introduces the Human-Centric Large Language Model (HCLLM) framework, which for the first time deeply integrates natural language processing, human-computer interaction, and responsible AI methodologies throughout all stages—from system design and data collection to training, evaluation, and deployment. The framework harmonizes ethical, economic, and technical objectives, offering developers actionable, principle-based guidance. Its forward-looking applicability and practical utility are demonstrated through a case study situated in future workplace scenarios, thereby advancing LLM development toward a genuinely human-centered paradigm.

Ethical AIHuman-Centered AIHuman-Computer Interaction

Hot Scholars

PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
CL

Chenghua Lin

Professor of Natural Language Processing, University of Manchester
Natural language processingnatural language generationmachine learning
YW

Yuxia Wang

MBZUAI
Natural Language Processing