Score
Designs and implements analyses, audits, and taxonomies that quantify and characterize labeler/judge disagreement in annotated data, including per-class disagreement rates, statistical tests for concentration (e.g., at decision boundaries), clustering of disagreement cases, and categorization of error modes. Produces diagnostic reports that evaluate the effects of aggregation methods (such as majority vote), identify underlying causes of disagreement, and recommend targeted remediation or relabeling strategies for each disagreement category.
This work addresses the challenge of substantial annotator disagreement in subjective NLP tasks, where conventional majority voting discards valuable perspective diversity and individual annotator modeling suffers from high cost and poor generalization. The authors propose an aggregation approach based on clustering annotators by consistency and systematically evaluate four strategies—majority voting, ensemble methods, multi-label learning, and multi-task learning—across three tasks (sentiment analysis, emotion classification, and hate speech detection) and 40 multilingual datasets. Experimental results demonstrate that integrating consistency-based annotator clustering with multi-label or multi-task learning effectively preserves annotation diversity while significantly improving classification performance, outperforming both majority voting and per-annotator modeling by leveraging disagreement as informative signal rather than noise.
This work addresses the challenge of annotation disagreement in subjective NLP tasks, where ambiguity in labeling criteria or overlapping category boundaries often leads to inconsistent judgments. The authors propose a pattern-level diagnostic framework that introduces an interpretable, criterion-level auditing mechanism prior to label aggregation. By collecting fine-grained evaluations from multiple annotators on individual labeling guidelines, the method systematically identifies two failure modes: unstable annotation standards and systematic category overlap. This approach provides the first structured attribution of disagreement early in the annotation pipeline, offering empirical grounding for refining annotation schemes. Evaluated on a commercial document task involving persuasion-value extraction, the analysis reveals that disagreements concentrate around a few unstable criteria, with nearly half of the sentences activating multiple categories. The diagnostic outcomes show strong alignment with domain expert assessments and effectively inform revisions to the annotation guidelines.
This study addresses the common yet often overlooked issue of subjective disagreement in multi-label sentiment annotation, which traditional approaches typically treat as noise and discard along with its underlying structural information. To better capture annotator uncertainty, the work proposes replacing hard majority voting with soft labels derived from vote proportions and intensity-weighted confidence, and introduces Soft Bernoulli Cross-Entropy (SoftBCE) for soft-supervised model training. Additionally, it incorporates a probabilistic alignment metric for evaluation and a data-driven diagnostic framework to analyze annotation discrepancies. Experimental results show that while hard labels yield marginally higher F1 scores, soft labels more faithfully represent the inherent uncertainty among annotators. This research establishes a novel paradigm and offers practical guidance for label aggregation, model training, and evaluation in multi-label sentiment analysis.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
This study investigates disagreement dynamics and their implications in collaborative Wikidata knowledge graph construction. Methodologically, it employs a mixed-methods approach: quantitative analysis—including temporal modeling and participation metrics—to characterize controversy evolution, and qualitative coding—encompassing thematic classification, role annotation, and discourse analysis—to identify disagreement types, interaction patterns, and participant roles. Key contributions include the first systematic quantification of Wikidata discussion practices: attribute deletion decisions take significantly longer than creation; over 50% of controversies remain unresolved; rational edits yield high-value insights yet exhibit low retention, with 25% of deeply engaged editors departing shortly thereafter; and overt conflict or vandalism is rare, reflecting an inclusive, consensus-oriented community culture. The findings reveal critical governance challenges—inefficient consensus formation and severe editor attrition—providing empirical grounding for improving collaborative mechanisms and discussion tools in knowledge graph curation.
Current evaluation methods struggle to uncover substantial disagreements among large language models (LLMs) in public opinion classification, potentially misleading policy decisions. This work proposes an interpretability-focused auditing framework that treats inter-model disagreement as a signal of semantic complexity, directing human review toward genuinely ambiguous opinions. Through multi-model comparisons, expert-defined scoring rules, and a two-stage annotation experiment, the study finds that thematic disagreements across models significantly outweigh variations caused by prompt perturbations within a single model. While expert rules mitigate superficial discrepancies, they fail to resolve deeper cognitive divergences. Moreover, human annotators frequently introduce novel interpretive frameworks absent from model outputs. Moving beyond conventional accuracy metrics, this paradigm highlights the diagnostic value of disagreement in interpretive coding for nuanced opinion analysis.
This study addresses a critical gap in machine learning education: the overreliance on pre-labeled datasets, which often obscures the subjectivity and ambiguity inherent in data annotation, leading students to place undue trust in model outputs. To counter this, the authors introduce an innovative pedagogical intervention that transforms manual annotation into an active learning tool. Students annotated hair coverage in skin lesion images using a three-point scale, followed by structured reflections via questionnaires. A cross-institutional experiment involving 43 participants from Fontys University of Applied Sciences (Netherlands) and the IT University of Copenhagen (Denmark) demonstrated that this approach significantly enhanced learners’ awareness of annotation ambiguity, dataset biases, and model limitations. Most participants acknowledged the influence of personal interpretation on labeling decisions and reported higher engagement compared to traditional instruction. This work provides the first empirical evidence supporting subjective annotation as an effective strategy for cultivating critical thinking about AI systems.
This study addresses the inconsistent metric selection and inadequate reporting practices in LLM-as-judge research, which hinder reproducibility and cross-study comparison. Through a systematic analysis of agreement measures between large language models and human evaluators, the work reveals mathematical equivalences among multiple correlation coefficients—such as Pearson, Spearman, phi, and Matthews—under binary scoring, clarifies the distinct utility of Cohen’s κ, and elucidates how handling abstentions fundamentally affects evaluation outcomes. Building on these statistical insights, the paper proposes a standardized reporting checklist that encompasses rating scales, treatment of abstentions and ties, coverage, confusion matrices, and aggregation levels. This framework substantially enhances the transparency, comparability, and reproducibility of LLM-as-judge evaluations.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study addresses the limitations of majority voting in annotating hate and offensive speech, which obscures substantial annotator disagreements on subjective boundaries and leads models to treat contested judgments as objective truths. Focusing on the HateXplain dataset, the authors systematically analyze disagreement patterns and evaluate three modeling approaches: hard-label BERT, soft-label models, and per-annotator multi-head architectures. Through chi-square tests and confidence analysis, they demonstrate that all models suffer a 22–28 percentage point accuracy drop on contentious samples, with the multi-head model achieving only 0.245 accuracy on offense-related disagreements. Critically, standard evaluation metrics fail to capture this degradation, as models often exhibit spuriously high confidence in incorrect predictions. The work exposes the structural bias inherent in majority voting for sensitive content annotation and its detrimental impact on model evaluation reliability.