harmonize annotations

Designs and implements mapping, reconciliation, and normalization procedures that align and transfer annotations across divergent schemes, vocabularies, and representation topologies, including projection or transfer between formats and explicit computation of inter-annotator agreement. Analyzes and documents annotation decisions and motivations, flags transformations that affect fidelity, and adapts normalization rules to accommodate genre- or domain-related variation.

harmonizeannotations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel approach to address the challenge novice modelers often face in ensuring semantic alignment between domain models and textual specifications during early software engineering phases. The method first employs natural language processing to preprocess specification texts and generates human-authored natural language descriptions for each model element. It then leverages a large language model (LLM) to compare these descriptions against the original specifications, automatically classifying their alignment status as aligned, misaligned, or uncertain, while providing interpretable evidence for each judgment. By uniquely integrating LLM capabilities with human-crafted model descriptions, the approach achieves high-precision semantic alignment verification, demonstrating near-perfect precision (≈100%) and 78% recall across multiple domain datasets. Individual element analysis requires between 18 seconds and one minute, indicating strong potential for integration into modeling tools.

Domain ModelModel ValidationSemantic Alignment

In high-quality machine translation (MT) human evaluation, performance gains are often obscured by annotation noise. To address this, we propose a two-stage MQM-based collaborative re-annotation method: building upon initial annotations, it introduces iterative human review and collaborative refinement to jointly optimize primary annotations, peer annotations, and automatic predictions. This work is the first to integrate collaborative editing behavior modeling into the MQM framework. It significantly improves annotation consistency (+18.3%) and error detection rate (+24.7%), effectively recovering errors missed in the first round. Experiments demonstrate that re-annotation substantially enhances the reliability and stability of evaluation outcomes, enabling more accurate reflection of model quality improvements in assessment scores. The proposed approach establishes a scalable, reproducible paradigm for high-precision MT evaluation.

Developing a two-stage MQM re-annotation techniqueEnhancing annotation quality by finding missed errorsImproving human evaluation methods for machine translation

Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study

Oct 13, 2025
KW
Kon Woo Kim
🏛️ The Graduate University for Advanced Studies, SOKENDAI | National Institute of Informatics | National Library of Medicine | Joint Support-Center for Data Science Research | Japanese-French Laboratory of Informatics | CNRS | Nantes University

This work addresses the challenge that human-authored annotation guidelines are poorly suited for large language model (LLM)-based text annotation due to their informal, ambiguous, and context-dependent nature. We propose a guideline refactoring method oriented toward LLM auditing: automatically transforming natural-language guidelines into structured, semantically precise, instruction-style rules aligned with LLM comprehension preferences. Our approach preserves original semantic intent while systematically enhancing executability and robustness. Evaluated on disease entity recognition using the NCBI Disease Corpus, the refactored guidelines significantly improve LLM annotation accuracy and inter-annotator consistency, enabling automated iterative guideline refinement. Empirical analysis further identifies critical failure modes—including instruction ambiguity and insufficient coverage of edge cases. This study establishes a novel paradigm and reusable methodological framework for building high-quality, LLM-native annotation infrastructure.

Addressing practical challenges in automated annotation workflow scalingRepurposing human annotation guidelines for LLM text annotationTransforming guidelines into explicit instructions for language models

This work addresses the challenge that large language models (LLMs) struggle to adhere to domain-specific gold-standard annotation guidelines in zero-shot settings. To mitigate this limitation, the authors propose a mediation framework that iteratively reuses and refines annotation guidelines, introducing guideline evolution as a novel alignment mechanism to enhance annotation consistency and accuracy under low-supervision conditions. The approach integrates reasoning-optimized variants from three major LLM families—GPT, Gemini, and DeepSeek—and employs iterative guideline consolidation and fine-tuning. Evaluated on biomedical named entity recognition benchmarks including NCBI Disease, BC5CDR, and BioRED, the method demonstrates significant improvements in the models’ compliance with expert annotation standards.

annotation guidelinesgold-standard benchmarksLarge Language Models

Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement

Jan 25, 2023
GA
Gavin Abercrombie
🏛️ Heriot-Watt University | Alana AI | Bocconi University

NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.

Assessing annotator inconsistency across multiple NLP classification tasksInvestigating reasons for annotator disagreement through quality control measuresMeasuring intra-annotator agreement for label stability in NLP tasks

Latest Papers

What's happening recently
View more

This study addresses the high cost and reliability challenges of employing large language models (LLMs) for structured annotation by proposing a dual-annotation-stream framework based on character-level alignment. The method automatically resolves unambiguous cases while routing conflicts to a browser-based interface for human adjudication. By integrating offline auditability, explicit logging strategies, and document- and span-level consistency computation, the framework supports field-level hybrid construction and direct export in original formats. Evaluated on a humanitarian benchmark, the system autonomously merges 8% of documents and precisely identifies 3,131 conflicts, substantially enhancing both review efficiency and result trustworthiness in human–machine collaborative annotation workflows.

annotation reconciliationdata auditingLLM extraction

Existing formality transfer datasets, such as GYAFC, frame the task as a binary symmetric transformation, leading models to generate “pseudo-formal” text that misaligns with human perceptions of absolute formality. To address this limitation, this work proposes a three-level formality framework—informal, casual, and formal—with “casual” serving as an intermediate anchor, and introduces 3LF, the first parallel dataset enabling continuous-spectrum formality transfer. Through a formal evaluation protocol, fine-tuning experiments across multiple models, and joint analysis using both human judgments and automatic metrics (e.g., F1), the study demonstrates the critical role of supervised structure in achieving style alignment. Training on 3LF substantially improves transfer performance: GPT-4.1-nano’s F1 score on informal-to-formal conversion rises from 0.06 to 0.88, a gain not replicable via in-context learning alone.

benchmark designcontrollable text generationformality transfer

This work addresses the lack of systematic alignment between large language models and human reviewers in survey evaluation, as well as the absence of a multidimensional, quantifiable assessment framework. To bridge this gap, the authors introduce SurveyReview—the first benchmark specifically designed for survey reviewing—comprising 675 survey papers and 1,630 structured review reports, along with standardized data splits and evaluation protocols. Building upon Qwen3-32B with LoRA fine-tuning and external knowledge augmentation, the proposed strong baseline model, SurveyAlign, translates free-form reviews into scores and justifications across four dimensions: readability, criticality, comprehensiveness, and structure. Experimental results demonstrate that SurveyAlign significantly outperforms GPT-5.2 with prompt-based evaluation, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 on the test set, thereby substantially improving alignment with human reviewers.

benchmarklarge language modelspeer review

This work addresses the high cost and substantial cognitive load associated with structured annotation, which hinder efficient allocation of labeling efforts among heterogeneous annotators such as humans and models. The authors propose a center-theory-based task decomposition approach that identifies semantic centers to constrain the output space, formally models reasoning load, and introduces an algorithm for allocating heterogeneous annotation resources. Notably, this is the first method to integrate task decomposition with annotator capability matching. Experimental results demonstrate that the proposed approach significantly reduces cognitive burden while simultaneously improving both annotation quality and cost efficiency under a fixed budget.

Annotation EfficiencyHeterogeneous AnnotatorsInferential Load

This work addresses systematic limitations in existing creative quality alignment (CQA) datasets, particularly their inadequate modeling of audience preferences and insufficient coverage of real-world logical constraints. To overcome these issues under stringent engineering and data scarcity conditions, the authors propose a low-resource CQA approach that leverages only around one hundred expert-annotated chain-of-thought (CoT) examples. By uncovering a dual mechanism between appreciation and generation tasks within conditional generative architectures, the method enables automatic transfer of calibrated knowledge from the appreciation module to the generation module. Experimental results demonstrate that the proposed framework substantially mitigates the shortcomings of current datasets and validates the practical feasibility of aligning generative models with nuanced creative quality metrics in real-world engineering settings.

Alignment Dataset BiasCalibrated SurpriseChain-of-Thought Fine-Tuning

Hot Scholars

ZM

Zdravko Marinov

Ph.D. Student at Karlrsruhe Institute of Technology
interactive segmentationmedical image analysisaction recognitiondomain adaptation
RK

Ranjay Krishna

University of Washington, Allen Institute for AI
Computer VisionNatural Language ProcessingMachine LearningHuman Computer Interaction
DD

Dorottya Demszky

Assistant Professor, Stanford University
natural language processingeducation data scienceteacher professional learning
LR

Lisa Raithel

BIFOLD, TU Berlin, DFKI
Information ExtractionBioNLPclinical decision supportmultilinguality
RR

Roland Roller

German Research Center for Artificial Intelligence (DFKI)
Natural Language ProcessingMedical NLPClinical Decision SupportAnonymization