Score
Design, build, and evaluate annotation systems and workflows that integrate human annotators and LLMs to produce structured labels, metadata, and human-editable prelabels; this includes schema-constrained and prompt-driven preannotation, calibrated and zero-shot LLM labeling, bootstrapping pipelines, and multi-stage diagnostic traces or candidate explanations for human review. Also develop and analyze crowdsourcing and annotator-management processes, validation and quality-control protocols, interfaces for hybrid human–LLM collaboration, and empirical studies and metrics that quantify accuracy, agreement, speed-up, and per-item labeling effort.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.
In management research, unstructured text annotation has long relied on crowdsourced human labor, while large language models (LLMs) offer efficiency and cost advantages but lack a systematic, reproducible framework for evaluating their applicability. Method: We propose SILICON, an LLM-powered text annotation workflow tailored to management research, integrating structured annotation guideline design, expert-derived baseline construction, iterative prompt optimization, and multi-model cross-validation—introducing, for the first time, a regression-based method for comparing LLM outputs. Contribution/Results: Validated via Krippendorff’s α reliability analysis and seven empirical case studies, SILICON demonstrates high agreement between LLM and expert annotations in single-label tasks (α > 0.8), but markedly reduced consistency in multi-label classification. Results confirm that expert baselines outperform crowdsourced annotations and that multi-model evaluation is indispensable. We publicly release a comprehensive practice guide and end-to-end implementation code, addressing a critical methodological gap in LLM-assisted qualitative research.
For subjective, multi-dimensional annotation tasks—such as search query clarification—current large language models (LLMs) still fall short of human-level performance in automated labeling, necessitating robust human-in-the-loop mechanisms. Method: We propose a lightweight human-in-the-loop annotation framework that leverages multi-LLM ensemble reasoning and confidence calibration to dynamically identify low-confidence and inter-model disagreement samples, thereby triggering targeted human review. This enables the construction of high-quality, multi-dimensional labeled datasets through a quality-controllable hybrid annotation pipeline. Contribution/Results: Experiments demonstrate that our approach maintains annotation consistency and reliability while reducing human effort by up to 45%. It significantly improves annotation efficiency and scalability, offering a cost-effective, robust paradigm for deploying LLMs in complex evaluation scenarios.
This study investigates whether large language models (LLMs) can reliably replace human annotators to mitigate the high cost and logistical complexity of human-subject studies in software engineering innovation evaluation. We systematically evaluate six state-of-the-art LLMs across ten code-related annotation tasks—including code summary quality assessment and defect repair judgment—using five public datasets. Methodologically, we propose *inter-model agreement* as a novel task-adaptability predictor and integrate confidence-threshold filtering to identify samples safe for LLM-only annotation, thereby establishing a hybrid human–LLM evaluation paradigm. Results show that LLMs achieve or approach human inter-annotator agreement (Krippendorff’s α ≥ 0.8) on multiple tasks; inter-model agreement strongly predicts task feasibility (AUC = 0.92); and confidence-based filtering raises replacement accuracy to 94.3%.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.
This study addresses the limitations of large language models (LLMs) in annotating complex social science constructs—such as climate mitigation pessimism—where autonomous labeling often yields suboptimal quality. To overcome this, the authors propose AnnotateThis, a human-centered interactive annotation system that introduces an innovative “LLM grounding” paradigm, deeply integrating expert knowledge into the LLM annotation pipeline. The system enables iterative co-evolution of conceptual definitions and model refinement through human–AI collaboration, interactive visualizations, and dynamic prompt optimization, functioning effectively both with and without ground-truth labels. Empirical evaluation demonstrates that, in labeled settings, AnnotateThis achieves a 0.15 improvement in F-Measure and a 0.23 gain in accuracy, significantly outperforming existing fully automated approaches.
This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.
This work addresses systematic biases, low agreement with human experts, irreproducibility, and data privacy risks inherent in large language models for automated annotation by proposing a highly aligned, deterministic, and open-source labeling framework based on a 4-bit quantized 1.7B-parameter small language model. Through task-aligned fine-tuning, a multidimensional scoring mechanism, and strategies involving data augmentation and regularization, the framework achieves substantially improved annotation quality with only limited human-annotated data. Experimental results demonstrate that the method outperforms the current best closed-source large models by 0.23 in Krippendorff’s α agreement metric and exhibits strong generalization across tasks such as sentiment classification. Notably, this study provides the first evidence that a quantized small model, when properly aligned via fine-tuning, can surpass state-of-the-art closed-source models in annotation consistency.
Current AI evaluation practices relying on human judgment are susceptible to anchoring effects and lack scalability, limiting their ability to provide reliable quality signals. This work proposes a human–AI collaborative evaluation framework in which humans focus on identifying salient information units—referred to as “nuggets”—and making value judgments, while large language models (LLMs) efficiently match model outputs against these nuggets. The approach integrates human oversight and automated scoring through an interactive annotation tool, a three-stage workflow, and an exportable nugget repository. By structuring human input around discrete, reusable semantic units, the framework substantially improves evaluation consistency, scalability, and accountability, thereby enhancing the reliability of LLM-as-a-Judge paradigms.
This work addresses the prevailing limitation in large language model (LLM) development, wherein human values are typically incorporated only post-training, lacking systematic integration across the model’s entire lifecycle. To bridge this gap, the paper introduces the Human-Centric Large Language Model (HCLLM) framework, which for the first time deeply integrates natural language processing, human-computer interaction, and responsible AI methodologies throughout all stages—from system design and data collection to training, evaluation, and deployment. The framework harmonizes ethical, economic, and technical objectives, offering developers actionable, principle-based guidance. Its forward-looking applicability and practical utility are demonstrated through a case study situated in future workplace scenarios, thereby advancing LLM development toward a genuinely human-centered paradigm.