Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional audio emotion recognition in precisely localizing emotional temporal boundaries within multi-speaker scenarios. To this end, we reformulate the task as temporal emotion localization by leveraging large language models that integrate masked language modeling with timestamp prediction losses to achieve fine-grained emotion localization. The core innovation lies in the proposed M-TAG supervision objective, which employs attention masking to mitigate shortcut dependencies and introduces distance-aware weighting to optimize temporal errors. Experimental results demonstrate that our approach significantly outperforms baseline models such as Flamingo-Next in both emotion recognition and temporal localization. These findings reveal the limitations of existing methods and establish a new paradigm for multi-speaker emotion analysis.
📝 Abstract
Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with temporally bounded speech spans and vocal tone descriptions. To support this formulation, we curate temporally annotated versions of existing emotion recognition datasets and construct recordings containing two to four affective speech spans, including overlapping speech. Training in this longer format with a standard language-modeling objective can degrade both emotion recognition and temporal grounding performance, while tone descriptions can provide shortcuts for emotion prediction. To address these challenges, we introduce Masked Temporal Affective Grounding (M-TAG), a supervised training objective that combines full-sequence language modeling with emotion and timestamp cross-entropy losses under attention masking. The masking varies the context visible to emotion-label tokens to reduce reliance on shortcuts and improve generalization, while the timestamp loss incorporates a distance-aware weight to penalize larger temporal errors. We evaluate EMO-TAG, a model fine-tuned using our dataset and objective, on emotion recognition and affective temporal-grounding against three AER and audio-language baselines: Flamingo-Next, Audio-Reasoner, and AffectGPT. Our results show that existing models achieve limited affective temporal-grounding despite competitive emotion recognition performance.
Problem

Research questions and friction points this paper is trying to address.

Audio Emotion Recognition
Temporal Affective Grounding
Multi-speaker
Overlapping Speech
Emotion Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Affective Grounding
Masked Temporal Affective Grounding
Audio Emotion Recognition
Attention Masking
Distance-aware Loss
🔎 Similar Papers
No similar papers found.