🤖 AI Summary
This work addresses the limitations of traditional speech emotion recognition (SER), which relies on hard labels and overlooks inter-annotator perceptual disagreement, thereby failing to capture the inherent uncertainty in human emotion judgments. To overcome this, the authors propose a distribution-based supervision approach, training a WavLM-Base multi-task model on the MSP-Podcast 2.0 dataset using soft labels derived from both primary annotator ratings and aggregated majority-minority voting. An entropy-aware curriculum strategy is introduced to prioritize learning from highly ambiguous samples. Evaluation via Jensen–Shannon divergence and Kullback–Leibler divergence demonstrates that the proposed method significantly reduces the discrepancy between model-predicted distributions and human vote distributions, particularly excelling on high-entropy samples. This study advances SER toward a soft-target paradigm that better reflects the uncertainty intrinsic to human emotional perception.
📝 Abstract
Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement. We study distribution-based supervision for 9-class SER on MSP-Podcast 2.0 using a WavLM-Base multitask model for categorical emotion and dimensional VAD. Hard-label training is compared with targets from primary and merged primary--secondary annotator vote distributions. Distributional objectives improve alignment with human vote distributions, reducing JSD/KLD relative to hard-label training. Analysis shows that hard supervision partly benefits from assigning ambiguous utterances to the residual Other class, whereas distributional supervision redistributes uncertainty across emotion categories. Entropy-stratified evaluation shows that high-ambiguity utterances remain challenging, but distribution-based supervision better captures perceptual uncertainty. These findings support moving beyond hard labels toward targets that reflect listener disagreement.