EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification

📅 2025-05-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In continuous-dimensional Valence-Arousal-Dominance (VAD) regression for Speech Emotion Recognition (SER), performance is hindered by annotation ambiguity and cross-corpus distribution shifts. To address this, we propose a spherical-guided structured joint learning framework: (1) VAD vectors are mapped onto the unit sphere, partitioned into multiple regions, and coupled with an auxiliary region classification task to impose discrete priors on continuous regression; (2) a style-aware dynamic weighted multi-head self-attention pooling layer is designed to enable fine-grained emotion-focused representation learning while disentangling speaker-specific stylistic variations. This work introduces the first paradigm integrating spherical coordinate transformation with region classification to jointly guide VAD regression. Evaluated on RAVDESS and IEMOCAP, our method reduces mean VAD regression error by 12.7% and significantly improves cross-corpus generalization and prediction stability—offering a more interpretable and robust approach to continuous-dimensional affective modeling.

Technology Category

Cognitive Modeling & Cognitive Systems: Affective ComputingNatural Language Processing: SpeechMachine Learning: Transfer, Domain Adaptation, Multi-Task Learning

Application Category

Search and Retrieval-Augmented AI: Vertical and domain-specific searchUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model that integrates spherical VAD region classification to guide VAD regression for improved emotion prediction. In our framework, VAD values are transformed into spherical coordinates that are divided into multiple spherical regions, and an auxiliary classification task predicts which spherical region each point belongs to, guiding the regression process. Additionally, we incorporate a dynamic weighting scheme and a style pooling layer with multi-head self-attention to capture spectral and temporal dynamics, further boosting performance. This combined training strategy reinforces structured learning and improves prediction consistency. Experimental results show that our approach exceeds baseline methods, confirming the validity of the proposed framework.
Problem

Research questions and friction points this paper is trying to address.

Enhancing speech emotion recognition using spherical VAD representation
Improving emotion prediction via auxiliary spherical region classification
Boosting performance with dynamic weighting and multi-head attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spherical VAD region classification for emotion prediction
Dynamic weighting scheme with multi-head self-attention
Joint training strategy for structured learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Deok-Hyeon Cho
Department of Artificial Intelligence, Korea University, Seoul, Korea
Hyung-Seok Oh
Hyung-Seok Oh
Korea Unviersity
Speech synthesis Deep Learning
Seung-Bin Kim
Seung-Bin Kim
Department of Artificial Intelligence, Korea University, Seoul, Korea
Speech Synthesis
S
Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Seoul, Korea