Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

📅 2026-06-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional speech emotion recognition (SER), which relies on hard labels and overlooks inter-annotator perceptual disagreement, thereby failing to capture the inherent uncertainty in human emotion judgments. To overcome this, the authors propose a distribution-based supervision approach, training a WavLM-Base multi-task model on the MSP-Podcast 2.0 dataset using soft labels derived from both primary annotator ratings and aggregated majority-minority voting. An entropy-aware curriculum strategy is introduced to prioritize learning from highly ambiguous samples. Evaluation via Jensen–Shannon divergence and Kullback–Leibler divergence demonstrates that the proposed method significantly reduces the discrepancy between model-predicted distributions and human vote distributions, particularly excelling on high-entropy samples. This study advances SER toward a soft-target paradigm that better reflects the uncertainty intrinsic to human emotional perception.
📝 Abstract
Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement. We study distribution-based supervision for 9-class SER on MSP-Podcast 2.0 using a WavLM-Base multitask model for categorical emotion and dimensional VAD. Hard-label training is compared with targets from primary and merged primary--secondary annotator vote distributions. Distributional objectives improve alignment with human vote distributions, reducing JSD/KLD relative to hard-label training. Analysis shows that hard supervision partly benefits from assigning ambiguous utterances to the residual Other class, whereas distributional supervision redistributes uncertainty across emotion categories. Entropy-stratified evaluation shows that high-ambiguity utterances remain challenging, but distribution-based supervision better captures perceptual uncertainty. These findings support moving beyond hard labels toward targets that reflect listener disagreement.
Problem

Research questions and friction points this paper is trying to address.

speech emotion recognition
annotation uncertainty
label ambiguity
distributional supervision
annotator disagreement
Innovation

Methods, ideas, or system contributions that make the work stand out.

annotation uncertainty
distributional supervision
entropy-aware curriculum
speech emotion recognition
soft labels
💼 Related Jobs
No related jobs found.
Z
Zahra Omidi
Center for Robust Speech Systems, The University of Texas at Dallas, USA
J
John H. L. Hansen
Center for Robust Speech Systems, The University of Texas at Dallas, USA