Human Feedback Driven Dynamic Speech Emotion Recognition

📅 2025-08-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Speech emotion is inherently time-varying, yet conventional methods often assume a static, single-label emotion per utterance—limiting their applicability to natural, real-time affective animation of 3D virtual agents. To address this, we propose a multi-stage training and human-feedback-driven optimization framework for dynamic speech emotion recognition. First, we model emotional mixtures using a Dirichlet distribution, enabling fine-grained, continuous emotion sequence prediction. Second, we integrate human feedback into a reinforcement learning loop to iteratively refine the model while substantially reducing reliance on dense manual annotations. Experiments demonstrate that Dirichlet-based modeling significantly outperforms sliding-window baselines, achieving a 4.2% F1-score gain on RAVDESS and comparable datasets. Moreover, annotation efficiency improves by approximately 60%. This work establishes a novel, interpretable, and optimization-friendly paradigm for dynamic emotion modeling, directly supporting expressive, real-time emotional animation in virtual humans.

Technology Category

Natural Language Processing: SpeechCognitive Modeling & Cognitive Systems: Affective ComputingHumans and AI: Emotional Intelligence

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study particularly focuses on the animation of emotional 3D avatars. We propose a multi-stage method that includes the training of a classical speech emotion recognition model, synthetic generation of emotional sequences, and further model improvement based on human feedback. Additionally, we introduce a novel approach to modeling emotional mixtures based on the Dirichlet distribution. The models are evaluated based on ground-truth emotions extracted from a dataset of 3D facial animations. We compare our models against the sliding window approach. Our experimental results show the effectiveness of Dirichlet-based approach in modeling emotional mixtures. Incorporating human feedback further improves the model quality while providing a simplified annotation procedure.
Problem

Research questions and friction points this paper is trying to address.

Dynamic speech emotion recognition with sequential emotional labels
Modeling emotional mixtures using Dirichlet distribution approach
Improving emotion recognition through human feedback integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-stage method combining SER training and synthetic generation
Novel emotional mixture modeling using Dirichlet distribution approach
Human feedback integration for model improvement and simplified annotation
🔎 Similar Papers
No similar papers found.
I
Ilya Fedorov
NVIDIA, Switzerland
D
Dmitry Korobchenko
NVIDIA, UK