Score
Designs and implements mapping, reconciliation, and normalization procedures that align and transfer annotations across divergent schemes, vocabularies, and representation topologies, including projection or transfer between formats and explicit computation of inter-annotator agreement. Analyzes and documents annotation decisions and motivations, flags transformations that affect fidelity, and adapts normalization rules to accommodate genre- or domain-related variation.
This work proposes a novel approach to address the challenge novice modelers often face in ensuring semantic alignment between domain models and textual specifications during early software engineering phases. The method first employs natural language processing to preprocess specification texts and generates human-authored natural language descriptions for each model element. It then leverages a large language model (LLM) to compare these descriptions against the original specifications, automatically classifying their alignment status as aligned, misaligned, or uncertain, while providing interpretable evidence for each judgment. By uniquely integrating LLM capabilities with human-crafted model descriptions, the approach achieves high-precision semantic alignment verification, demonstrating near-perfect precision (≈100%) and 78% recall across multiple domain datasets. Individual element analysis requires between 18 seconds and one minute, indicating strong potential for integration into modeling tools.
In high-quality machine translation (MT) human evaluation, performance gains are often obscured by annotation noise. To address this, we propose a two-stage MQM-based collaborative re-annotation method: building upon initial annotations, it introduces iterative human review and collaborative refinement to jointly optimize primary annotations, peer annotations, and automatic predictions. This work is the first to integrate collaborative editing behavior modeling into the MQM framework. It significantly improves annotation consistency (+18.3%) and error detection rate (+24.7%), effectively recovering errors missed in the first round. Experiments demonstrate that re-annotation substantially enhances the reliability and stability of evaluation outcomes, enabling more accurate reflection of model quality improvements in assessment scores. The proposed approach establishes a scalable, reproducible paradigm for high-precision MT evaluation.
This work addresses the challenge that human-authored annotation guidelines are poorly suited for large language model (LLM)-based text annotation due to their informal, ambiguous, and context-dependent nature. We propose a guideline refactoring method oriented toward LLM auditing: automatically transforming natural-language guidelines into structured, semantically precise, instruction-style rules aligned with LLM comprehension preferences. Our approach preserves original semantic intent while systematically enhancing executability and robustness. Evaluated on disease entity recognition using the NCBI Disease Corpus, the refactored guidelines significantly improve LLM annotation accuracy and inter-annotator consistency, enabling automated iterative guideline refinement. Empirical analysis further identifies critical failure modes—including instruction ambiguity and insufficient coverage of edge cases. This study establishes a novel paradigm and reusable methodological framework for building high-quality, LLM-native annotation infrastructure.
This work addresses the challenge that large language models (LLMs) struggle to adhere to domain-specific gold-standard annotation guidelines in zero-shot settings. To mitigate this limitation, the authors propose a mediation framework that iteratively reuses and refines annotation guidelines, introducing guideline evolution as a novel alignment mechanism to enhance annotation consistency and accuracy under low-supervision conditions. The approach integrates reasoning-optimized variants from three major LLM families—GPT, Gemini, and DeepSeek—and employs iterative guideline consolidation and fine-tuning. Evaluated on biomedical named entity recognition benchmarks including NCBI Disease, BC5CDR, and BioRED, the method demonstrates significant improvements in the models’ compliance with expert annotation standards.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
This study addresses the high cost and reliability challenges of employing large language models (LLMs) for structured annotation by proposing a dual-annotation-stream framework based on character-level alignment. The method automatically resolves unambiguous cases while routing conflicts to a browser-based interface for human adjudication. By integrating offline auditability, explicit logging strategies, and document- and span-level consistency computation, the framework supports field-level hybrid construction and direct export in original formats. Evaluated on a humanitarian benchmark, the system autonomously merges 8% of documents and precisely identifies 3,131 conflicts, substantially enhancing both review efficiency and result trustworthiness in human–machine collaborative annotation workflows.
Existing formality transfer datasets, such as GYAFC, frame the task as a binary symmetric transformation, leading models to generate “pseudo-formal” text that misaligns with human perceptions of absolute formality. To address this limitation, this work proposes a three-level formality framework—informal, casual, and formal—with “casual” serving as an intermediate anchor, and introduces 3LF, the first parallel dataset enabling continuous-spectrum formality transfer. Through a formal evaluation protocol, fine-tuning experiments across multiple models, and joint analysis using both human judgments and automatic metrics (e.g., F1), the study demonstrates the critical role of supervised structure in achieving style alignment. Training on 3LF substantially improves transfer performance: GPT-4.1-nano’s F1 score on informal-to-formal conversion rises from 0.06 to 0.88, a gain not replicable via in-context learning alone.
This work addresses the lack of systematic alignment between large language models and human reviewers in survey evaluation, as well as the absence of a multidimensional, quantifiable assessment framework. To bridge this gap, the authors introduce SurveyReview—the first benchmark specifically designed for survey reviewing—comprising 675 survey papers and 1,630 structured review reports, along with standardized data splits and evaluation protocols. Building upon Qwen3-32B with LoRA fine-tuning and external knowledge augmentation, the proposed strong baseline model, SurveyAlign, translates free-form reviews into scores and justifications across four dimensions: readability, criticality, comprehensiveness, and structure. Experimental results demonstrate that SurveyAlign significantly outperforms GPT-5.2 with prompt-based evaluation, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 on the test set, thereby substantially improving alignment with human reviewers.
This work addresses the high cost and substantial cognitive load associated with structured annotation, which hinder efficient allocation of labeling efforts among heterogeneous annotators such as humans and models. The authors propose a center-theory-based task decomposition approach that identifies semantic centers to constrain the output space, formally models reasoning load, and introduces an algorithm for allocating heterogeneous annotation resources. Notably, this is the first method to integrate task decomposition with annotator capability matching. Experimental results demonstrate that the proposed approach significantly reduces cognitive burden while simultaneously improving both annotation quality and cost efficiency under a fixed budget.
This work addresses systematic limitations in existing creative quality alignment (CQA) datasets, particularly their inadequate modeling of audience preferences and insufficient coverage of real-world logical constraints. To overcome these issues under stringent engineering and data scarcity conditions, the authors propose a low-resource CQA approach that leverages only around one hundred expert-annotated chain-of-thought (CoT) examples. By uncovering a dual mechanism between appreciation and generation tasks within conditional generative architectures, the method enables automatic transfer of calibrated knowledge from the appreciation module to the generation module. Experimental results demonstrate that the proposed framework substantially mitigates the shortcomings of current datasets and validates the practical feasibility of aligning generative models with nuanced creative quality metrics in real-world engineering settings.