label distribution learning

Designs, builds, and evaluates models and pipelines that predict or estimate per-item probability distributions over labels (soft/probabilistic or vote-share labels) and methods to aggregate annotator votes into those distributions. Implements training procedures and distributional supervision (e.g., KL/JS-based losses, soft-target cross-entropy), weighting schemes such as intensity-weighted labels, and analyses of distributional mismatch, entropy correlation, and preservation of annotator disagreement.

labeldistributionlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses annotator disagreement by proposing hard-label strategies—such as multipass annotation and stochastic label sampling (SLS)—as alternatives to conventional soft-label training, thereby leveraging the full annotation distribution rather than treating it as noise. Theoretical analysis grounded in cross-entropy optimization reveals that hard-label methods converge to flatter loss basins under label sparsity, enhancing generalization. Empirical results demonstrate that these approaches significantly outperform soft-label methods on CIFAR-10H, match SLS performance under fully annotated settings, and exhibit superior out-of-distribution detection capabilities on SVHN and CIFAR-100.

annotation distributionannotator disagreementhard labels

This work addresses the pervasive annotator disagreement problem in crowdsourced labeling. We propose an enhanced perspectivist modeling approach built upon the DisCo neural architecture: (i) incorporating annotator metadata to enrich input representations; (ii) introducing a disagreement-aware loss reweighting mechanism; and (iii) jointly modeling sample-level soft label distributions and annotator-level perspective-specific evaluations. Unlike conventional methods assuming a single underlying label distribution, our framework explicitly captures the heterogeneity of annotation behaviors. Experiments on three benchmark datasets demonstrate substantial improvements in soft label prediction accuracy and perspective-aware evaluation performance. Notably, the method achieves robust gains across error distribution metrics (e.g., KL divergence, ECE) and calibration quality, validating its capacity to effectively model complex, multi-perspective labeling patterns.

Enhancing DisCo with metadata and loss reweighting for better predictionsImproving evaluation metrics through disagreement-aware modeling techniquesModeling annotator disagreement via soft label distribution prediction

This study addresses the common yet often overlooked issue of subjective disagreement in multi-label sentiment annotation, which traditional approaches typically treat as noise and discard along with its underlying structural information. To better capture annotator uncertainty, the work proposes replacing hard majority voting with soft labels derived from vote proportions and intensity-weighted confidence, and introduces Soft Bernoulli Cross-Entropy (SoftBCE) for soft-supervised model training. Additionally, it incorporates a probabilistic alignment metric for evaluation and a data-driven diagnostic framework to analyze annotation discrepancies. Experimental results show that while hard labels yield marginally higher F1 scores, soft labels more faithfully represent the inherent uncertainty among annotators. This research establishes a novel paradigm and offers practical guidance for label aggregation, model training, and evaluation in multi-label sentiment analysis.

annotator disagreementemotion annotationlabel aggregation

This study investigates how to allocate annotation resources effectively according to target evaluation metrics to capture annotator disagreement. Leveraging the ChaosNLI dataset and fine-tuning DeBERTa and RoBERTa models with soft labels, the work reveals for the first time that the annotation saturation point is highly metric-dependent: distribution matching saturates at around 10 annotators (achieving 87–95% of maximal performance gain), whereas disagreement detection requires 20–50 annotators. The experiments demonstrate that soft labels effectively distinguish ambiguous from unambiguous examples, achieving an entropy correlation of 0.643—significantly outperforming conventional label smoothing methods (0.45–0.49). These findings underscore the necessity and efficacy of adopting evaluation-metric-driven annotation strategies.

annotation saturationannotator disagreementevaluation metric

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

May 02, 2024
MP
Maja Pavlovic
🏛️ Queen Mary University of London | University of Utrecht

This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.

Addressing limitations like bias and prompt sensitivityComparing human and GPT-generated opinion distributionsEvaluating LLMs' effectiveness in data annotation tasks

Latest Papers

What's happening recently
View more

This work addresses the challenge of accurately recovering minority-class labels in crowdsourced annotation tasks under severe class imbalance. The authors propose a generative label aggregation model that jointly captures category-dependent item difficulty and annotator competence, departing from conventional assumptions in the field. By re-examining the Condorcet jury theorem under class-imbalanced settings, they theoretically demonstrate that majority voting asymptotically preserves the original class distribution. Empirical evaluation across 33 real-world multiclass crowdsourcing datasets shows that the proposed model substantially improves recall for minority classes while maintaining competitive overall accuracy.

class-dependent annotator accuracyimbalanced crowdsourcingitem difficulty

This work addresses the challenge of substantial annotator disagreement in subjective NLP tasks, where conventional majority voting discards valuable perspective diversity and individual annotator modeling suffers from high cost and poor generalization. The authors propose an aggregation approach based on clustering annotators by consistency and systematically evaluate four strategies—majority voting, ensemble methods, multi-label learning, and multi-task learning—across three tasks (sentiment analysis, emotion classification, and hate speech detection) and 40 multilingual datasets. Experimental results demonstrate that integrating consistency-based annotator clustering with multi-label or multi-task learning effectively preserves annotation diversity while significantly improving classification performance, outperforming both majority voting and per-annotator modeling by leveraging disagreement as informative signal rather than noise.

annotation disagreementannotator perspectiveslabel aggregation

This study addresses the issue of small-scale leaderboards in preference benchmarks obscuring annotation disagreements. Through comparative analysis of multiple annotation pools, bootstrap resampling, and LLM evaluator assessment, we quantify these discrepancies. Results reveal that 23.6% of entries exhibit divergence between expert and crowdsourced labels. Furthermore, the apparent stability of current rankings is illusory; as model counts increase to 10 and 20, rank permutation probabilities surge to 86% and 99.97%, respectively. Additionally, LLM judges demonstrate significant bias toward crowdsourcing standards. This work provides the first quantitative assessment of how annotation variance impacts rankings, exposing the misleading nature of small-sample consistency and group bias in automated evaluation, thereby challenging prevailing assumptions regarding benchmark reliability.

Annotator DisagreementInter-annotator VariabilityLeaderboard Robustness

This work addresses the limitations of current large language models in text-based regression tasks, which struggle to accurately model the full conditional distribution and lack direct input-output associations as well as local contextual support. To overcome these challenges, the authors propose an end-to-end distribution prediction framework based on quantile tokens. Specifically, they introduce dedicated quantile tokens that establish direct pathways between inputs and quantile outputs via self-attention mechanisms, while incorporating retrieval-augmented empirical distributions from semantically similar neighbors for supervision. Theoretical analysis of the quantile regression loss function is provided, and experiments on the Inside Airbnb and StackSample datasets demonstrate significant improvements: compared to baselines, the method reduces MAPE by approximately 4 percentage points and narrows prediction intervals by a factor of two, with particularly notable gains in low-data or high-difficulty scenarios.

conditional distributiondistributional regressionempirical quantiles

This work addresses the limitations of traditional speech emotion recognition (SER), which relies on hard labels and overlooks inter-annotator perceptual disagreement, thereby failing to capture the inherent uncertainty in human emotion judgments. To overcome this, the authors propose a distribution-based supervision approach, training a WavLM-Base multi-task model on the MSP-Podcast 2.0 dataset using soft labels derived from both primary annotator ratings and aggregated majority-minority voting. An entropy-aware curriculum strategy is introduced to prioritize learning from highly ambiguous samples. Evaluation via Jensen–Shannon divergence and Kullback–Leibler divergence demonstrates that the proposed method significantly reduces the discrepancy between model-predicted distributions and human vote distributions, particularly excelling on high-entropy samples. This study advances SER toward a soft-target paradigm that better reflects the uncertainty intrinsic to human emotional perception.

annotation uncertaintyannotator disagreementdistributional supervision

Hot Scholars

FG

Fosca Giannotti

professor at Scuola Normale Superiore di Pisa
explainable artificial intelligencetrustworthy Aidata miningsocial network analysis
GG

Gizem Gezici

Scuola Normale Superiore, Pisa, ITALY
Natural Language ProcessingInformation RetrievalMachine Learning
BP

Barbara Plank

Professor, LMU Munich, Visiting Prof ITU Copenhagen
Natural Language ProcessingComputational LinguisticsMachine LearningTransfer Learning
MS

Masashi Sugiyama

Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence
JW

Jingyao Wu

MIT-Novo Nordisk AI Postdoctoral Fellow, MIT Media Lab
emotion recognitionaffective computingmachine learningspeech processing