instruction-guided annotation

Designs and writes precise, testable annotation instructions and label schemas, and implements instruction-driven labeling workflows for human annotators or automated labelers to produce consistent, high-quality labeled data. Defines edge cases, examples, and acceptance criteria and analyzes annotator adherence and inter-annotator agreement to iterate and refine the instructions.

instruction-guidedannotation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.

annotation qualitydata annotationmeasurement

This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.

annotation driftcategory definitionscontent moderation

Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop

Nov 07, 2024
EA
Ekaterina Artemova
🏛️ Toloka AI | Nebius AI | University of Stuttgart

High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.

Annotation CostMachine LearningNatural Language Processing

Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement

Jan 25, 2023
GA
Gavin Abercrombie
🏛️ Heriot-Watt University | Alana AI | Bocconi University

NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.

Assessing annotator inconsistency across multiple NLP classification tasksInvestigating reasons for annotator disagreement through quality control measuresMeasuring intra-annotator agreement for label stability in NLP tasks

Can LLMs Replace Manual Annotation of Software Engineering Artifacts?

Aug 10, 2024
TA
Toufique Ahmed
🏛️ University of California, Davis | Singapore Management University | University of Stuttgart

This study investigates whether large language models (LLMs) can reliably replace human annotators to mitigate the high cost and logistical complexity of human-subject studies in software engineering innovation evaluation. We systematically evaluate six state-of-the-art LLMs across ten code-related annotation tasks—including code summary quality assessment and defect repair judgment—using five public datasets. Methodologically, we propose *inter-model agreement* as a novel task-adaptability predictor and integrate confidence-threshold filtering to identify samples safe for LLM-only annotation, thereby establishing a hybrid human–LLM evaluation paradigm. Results show that LLMs achieve or approach human inter-annotator agreement (Krippendorff’s α ≥ 0.8) on multiple tasks; inter-model agreement strongly predicts task feasibility (AUC = 0.92); and confidence-based filtering raises replacement accuracy to 94.3%.

LLMs replace human annotationModel-model agreement predictorSoftware engineering artifact evaluation

Latest Papers

What's happening recently
View more

This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.

annotation reportingannotation validityhuman annotation

Current AI evaluation practices relying on human judgment are susceptible to anchoring effects and lack scalability, limiting their ability to provide reliable quality signals. This work proposes a human–AI collaborative evaluation framework in which humans focus on identifying salient information units—referred to as “nuggets”—and making value judgments, while large language models (LLMs) efficiently match model outputs against these nuggets. The approach integrates human oversight and automated scoring through an interactive annotation tool, a three-stage workflow, and an exportable nugget repository. By structuring human input around discrete, reusable semantic units, the framework substantially improves evaluation consistency, scalability, and accountability, thereby enhancing the reliability of LLM-as-a-Judge paradigms.

accountable evaluationhuman judgmentHuman-in-the-Loop

This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.

annotation pipelinesdata qualityerror propagation

This study investigates how large language models integrate instructional cues from system prompts, user prompts, and JSON schemas in structured output tasks, with particular focus on performance degradation when these sources conflict. Through single-field classification experiments across ten models from the GPT and Claude families, the authors conduct ablation studies on instruction placement, conflict scenarios, and interventions involving intermediate reasoning fields. Findings reveal that JSON schemas are not passive metadata but can substantially override prompt-based instructions: schema descriptions alone yield 13 percentage points lower accuracy than system prompts in non-conflicting settings and up to 45 points lower under conflict. Introducing intermediate reasoning fields improves accuracy by 15–24 points, even surpassing prompt-only approaches. These results motivate a unified view of prompts and schemas as a cohesive instruction interface.

instruction placementLLM behaviorprompt conflict

Hot Scholars

LY

Li Yuan

Research Associate, University of Science & Technology of China (USTC)
Antibiotic resistanceWastewater treatmentEnvironmental bioremediationAnaerobic digestion
AF

Aarash Feizi

PhD student in Computer Science, McGill University
Representation LearningSelf-Supervised LearningGraph Representation Learning
SN

Shravan Nayak

Mila
Vision and LanguageCultureGeo-diversityMultilinguality
XJ

Xiangru Jian

University of Waterloo
MultimodalityLLMGNNDatabase
SR

Sai Rajeswar

Staff Research Scientist, Adjunct Professor, Mila, ServiceNow
machine learninggenerative modelsreinforcement learning