Score
Design and train models that represent individual annotators' labeling behavior—e.g., per-annotator heads, latent embeddings, or explicit bias and variance parameters—to estimate and decompose each annotator's systematic bias and variability relative to a reference label. Use those annotator-specific representations to analyze how annotator variability propagates through predictions, enable controlled perturbation studies, measure annotator–model similarity (for example via cosine), and identify clusters of annotator tendencies.
Existing individual tendency learning (ITL) methods lack a unified, quantifiable framework to rigorously assess whether they genuinely capture annotator-level behavioral differences and yield behaviorally plausible explanations. This paper introduces the first evaluation framework for ITL in multi-annotator settings. Its core contributions are: (1) the Difference-aware Inter-annotator Consistency (DIC) metric, which quantifies a model’s ability to capture heterogeneity in annotator behavior; and (2) the Behavior-aligned Explainability (BAE) metric, the first to jointly evaluate the alignment between ITL-generated explanations and empirically observed annotator behavior. The framework integrates multidimensional scaling, predictive similarity structure comparison, and explanation verification, grounded in real-world annotation data. Extensive experiments demonstrate that DIC and BAE effectively distinguish state-of-the-art ITL methods in both tendency modeling fidelity and explanation plausibility, establishing a reliable, behaviorally grounded benchmark for future ITL research.
This study addresses the “quality–diversity trade-off” in subjective annotation tasks: conventional spam annotator filtering methods treat label variation as noise, erroneously removing reliable annotators holding minority opinions and thereby distorting the true opinion distribution. The authors argue that genuine spam annotators tend to exhibit *fixed-response behavior*—repeatedly selecting identical labels—not random guessing; thus, spam behavior must be redefined accordingly. They design multiple heuristic filtering strategies and systematically evaluate them on synthetically noisy data. Results show that annotation bias remains acceptable only when ≤5% of annotators are removed; most existing methods fail to detect fixed-response spammers and frequently misclassify non-random yet reliable dissenting annotators as noise. The core contribution is the insight that label diversity itself constitutes meaningful signal—not noise—in subjective tasks, and the proposal of *response fixity*, rather than inter-annotator agreement, as a principled criterion for spam detection.
This study addresses the lack of systematic understanding regarding how annotator characteristics and textual linguistic properties jointly influence annotation variability in harmful language detection. Integrating sociolinguistic features of annotators—including demographic attributes and attitudinal measures—with computational linguistic metrics of text, the authors conduct large-scale statistical modeling and joint analyses across four harmful language datasets. They uncover significant interaction effects between annotator and text features, demonstrating that lexical cues and annotator attitudes exert strong, interdependent influences on labeling outcomes. Notably, these interaction patterns vary substantially across datasets. These findings challenge prevailing practices that overlook such complexities and underscore the necessity of explicitly accounting for annotator–text interactions when developing and generalizing harmful language detection models.
In speech emotion recognition, the subjectivity and inter-annotator variability inherent in multi-annotator labels are often obscured by simplistic label averaging, leading to distorted modeling of emotional dynamics. To address this, we propose an end-to-end multitask framework that, for the first time, jointly predicts individual annotator identity and continuous emotion distributions—e.g., kernel density estimates or parametric distributions—during training. This explicitly models annotator behavioral heterogeneity while preserving population-level variability. Our approach eliminates label averaging and instead integrates annotator modeling directly into the distribution learning process, enabling co-optimization of annotator-specific characteristics and emotion distributions. Evaluated under both cross-corpus and in-corpus settings, our method achieves statistically significant improvements over state-of-the-art approaches in emotion distribution prediction accuracy. It more faithfully captures emotion subjectivity and annotator disagreement, offering a principled solution to modeling annotation uncertainty in affective computing.
Deep learning models often inherit implicit biases, and existing unsupervised debiasing methods rely on latent-space clustering to generate pseudo-labels—lacking semantic interpretability and hindering expert validation. To address this, we propose the first semantics-driven framework for diagnosing implicit bias: it requires neither human annotations nor prior assumptions about bias types. Our approach leverages text-guided latent-space disentanglement, task-relevance distillation, bias-semantic mapping, and a generative explanation module to automatically identify and semantically name non-representative bias features actually learned by the model. This enables explicit, interpretable bias identification and attribution, supporting both in-training intervention and post-hoc verification. Evaluated across multiple benchmarks, our method significantly improves bias detection accuracy and interpretability, demonstrating strong generalizability and practical operability.
This work proposes to model the stable signals underlying human label variation (HLV) by learning individualized labeling and reasoning behaviors from free-text explanations provided by human annotators. Addressing the limitation of existing approaches that overlook annotator-specific explanation styles, the authors introduce Cross-Annotator Preference Optimization (CAPO), which, for the first time, treats HLV as a learnable signal of individualized explanatory behavior. By disentangling input content effects, CAPO effectively captures annotator-specific reasoning patterns. Integrated with prompt engineering and supervised fine-tuning, the method significantly outperforms baselines on natural language inference and paraphrase judgment tasks, achieving higher accuracy in aggregated behavioral imitation while preserving individual reasoning styles and improving attribution quality under human evaluation.
This study addresses the challenges of quality assessment and contamination detection in large model training data, where absolute ground truth is often unavailable. We introduce the concept of "internal annotator dynamics," which analyzes a single annotator's self-reproducibility consistency on identical texts over time. Using the segmentation and causal-functional annotation of Sumerian mythology as an experimental testbed, we compare annotations produced by human experts and large language models (LLMs) across different time points. Our findings reveal that human annotators exhibit a distinctive pattern characterized by "stable boundaries but drifting labels." We demonstrate that this temporal signature can serve as an effective metric for data quality assessment and contamination detection, enabling the reliable identification of annotations that appear human-generated but are actually produced by LLMs.
This study addresses the long-standing misconception that estimator ranking inconsistencies in data attribution stem from approximation errors, revealing instead that they originate from counterfactual norm mismatches. By formalizing influence as a counterfactual estimator, this work establishes norm analysis as a necessary prerequisite for comparing estimators. It derives local decompositions to analytically characterize signal interaction mechanisms, validated through linearized approximations and controlled experiments. The primary contribution is the first demonstration that behavioral proxy selection critically impacts attribution quality, proving that differing norms directly induce ranking discrepancies. Furthermore, the proposed behavior-aligned norm successfully identifies target samples overlooked by default methods, substantially improving attribution accuracy.
This work proposes a context-aware multi-model optimization approach to overcome the limitations of relying on a single optimal model. By systematically identifying multiple models that exhibit substantially different feature selections yet achieve comparable predictive performance, the method preserves overall accuracy while enhancing interpretability. Applied to the METABRIC gene expression dataset, the approach successfully generates a diverse ensemble of high-performing models, uncovering multiple plausible biological interpretation pathways underlying the data. Compared to baseline methods, the resulting models demonstrate superior trade-offs between feature dissimilarity and performance consistency, thereby significantly improving both model interpretability and scientific insight.
This study addresses the common yet often overlooked issue of subjective disagreement in multi-label sentiment annotation, which traditional approaches typically treat as noise and discard along with its underlying structural information. To better capture annotator uncertainty, the work proposes replacing hard majority voting with soft labels derived from vote proportions and intensity-weighted confidence, and introduces Soft Bernoulli Cross-Entropy (SoftBCE) for soft-supervised model training. Additionally, it incorporates a probabilistic alignment metric for evaluation and a data-driven diagnostic framework to analyze annotation discrepancies. Experimental results show that while hard labels yield marginally higher F1 scores, soft labels more faithfully represent the inherent uncertainty among annotators. This research establishes a novel paradigm and offers practical guidance for label aggregation, model training, and evaluation in multi-label sentiment analysis.