Score
Designs and implements analyses and metrics to measure consistency and agreement among multiple raters—human annotators or automated systems—including pairwise and multi‑rater agreement statistics, inter‑rater reliability coefficients, and calibration checks. Performs diagnostic work to identify dimensions with weak agreement, investigate causes of disagreement, adjudicate conflicts, refine annotation guidelines, and train or calibrate raters and system outputs to improve reliability.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
Amid the growing diversity of natural language processing tasks, existing inter-annotator agreement (IAA) metrics often suffer from limited applicability and interpretability when confronted with heterogeneous task types, label imbalance, and missing data. This work systematically reviews the theoretical foundations and practical methodologies of IAA, offering the first structured integration of mainstream metrics—such as Cohen’s Kappa and Krippendorff’s Alpha—organized by task type. It clarifies their underlying assumptions and delineates their boundaries of applicability. Furthermore, the study proposes a reliability assessment strategy that combines confidence intervals with analysis of disagreement patterns. By providing a clear, principled guide for selecting IAA metrics, this research significantly enhances the transparency, reproducibility, and scientific rigor of human annotation and evaluation practices in the NLP community.
This study addresses the inconsistent metric selection and inadequate reporting practices in LLM-as-judge research, which hinder reproducibility and cross-study comparison. Through a systematic analysis of agreement measures between large language models and human evaluators, the work reveals mathematical equivalences among multiple correlation coefficients—such as Pearson, Spearman, phi, and Matthews—under binary scoring, clarifies the distinct utility of Cohen’s κ, and elucidates how handling abstentions fundamentally affects evaluation outcomes. Building on these statistical insights, the paper proposes a standardized reporting checklist that encompasses rating scales, treatment of abstentions and ties, coverage, confusion matrices, and aggregation levels. This framework substantially enhances the transparency, comparability, and reproducibility of LLM-as-judge evaluations.
Existing Cohen’s and Fleiss’ kappa statistics are restricted to single-label classification and lack robustness in realistic settings—such as comorbid psychiatric diagnoses or multi-behavior coding—where multi-label annotations, hierarchical category structures, variable numbers of annotators, and missing labels commonly occur. This paper introduces a generalized κ statistic: the first extension of Fleiss’ kappa to multi-label settings; it incorporates a category-weighting matrix and hierarchical distance metrics to capture semantic similarity among labels, and employs probabilistic expectation-based estimation to ensure robustness against missing data and variable annotator counts. Theoretically, it strictly reduces to classical Fleiss’ kappa under single-label conditions. Implemented in R and Excel, the method is validated through rigorous mathematical derivation and empirical application to psychiatric multi-diagnosis data. Its interpretable range is formally established as [−1, 1], with principled guidelines for interpretation.
This study addresses the reproducibility crisis in AI evaluation, which stems from subjective biases in human annotation and limited data leading to unstable results. The authors propose a multilevel bootstrapping framework that, for the first time, leverages large-scale human rating data—complete with persistent annotator identifiers—to model annotator behavior and systematically quantify the joint impact of the number of items (N) and the number of ratings per item (K) on evaluation reproducibility. Through multilevel bootstrap resampling, variance modeling, and significance analysis, the work uncovers the trade-offs inherent in N–K configurations and derives optimal strategies for reliable assessment design. This approach provides a methodological foundation for establishing robust, data-driven paradigms in AI evaluation.
This work addresses the lack of systematic alignment between large language models and human reviewers in survey evaluation, as well as the absence of a multidimensional, quantifiable assessment framework. To bridge this gap, the authors introduce SurveyReview—the first benchmark specifically designed for survey reviewing—comprising 675 survey papers and 1,630 structured review reports, along with standardized data splits and evaluation protocols. Building upon Qwen3-32B with LoRA fine-tuning and external knowledge augmentation, the proposed strong baseline model, SurveyAlign, translates free-form reviews into scores and justifications across four dimensions: readability, criticality, comprehensiveness, and structure. Experimental results demonstrate that SurveyAlign significantly outperforms GPT-5.2 with prompt-based evaluation, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 on the test set, thereby substantially improving alignment with human reviewers.
This study challenges the conventional reliance on intercoder agreement among humans as the gold standard for evaluating qualitative coding, highlighting that human consensus may embed systematic biases. To address this, the authors propose a transferable blind expert validation protocol and a code-level human–AI collaboration framework. They conduct a symmetric evaluation of coding outputs from five large language models (LLMs) and three human coders across 72 hierarchical codes applied to 2,560 educator messages. Integrating Jaccard similarity, Bradley–Terry ranking models, and blinded expert judgments, the findings reveal no significant preference for human over LLM-generated codes (48.5% vs. 51.5%), with two LLMs outperforming two human coders in expert rankings—thereby questioning the paradigm that equates human agreement with coding quality.
This study addresses the common yet often overlooked issue of subjective disagreement in multi-label sentiment annotation, which traditional approaches typically treat as noise and discard along with its underlying structural information. To better capture annotator uncertainty, the work proposes replacing hard majority voting with soft labels derived from vote proportions and intensity-weighted confidence, and introduces Soft Bernoulli Cross-Entropy (SoftBCE) for soft-supervised model training. Additionally, it incorporates a probabilistic alignment metric for evaluation and a data-driven diagnostic framework to analyze annotation discrepancies. Experimental results show that while hard labels yield marginally higher F1 scores, soft labels more faithfully represent the inherent uncertainty among annotators. This research establishes a novel paradigm and offers practical guidance for label aggregation, model training, and evaluation in multi-label sentiment analysis.
Existing tools for measuring political orientation perform poorly in non-Western, data-scarce regions. This work proposes an innovative approach that treats large language models as fallible expert raters, constructing a multi-model ensemble panel guided by axially defined ideological anchors, a mechanism separating stated beliefs from observed behaviors, and explicit rules distinguishing zero scores from missing values. Rater consistency is rigorously evaluated using Krippendorff’s alpha and triple-blind assessments, achieving high inter-rater agreement (α = 0.86) across nine models. The incorporation of formal definitions significantly reduces scoring discrepancies and enhances inter-model correlations. Disagreement analysis reveals that divergences in political stance judgments stem primarily from interpretive differences rather than random error, substantially improving the method’s cross-regional applicability.