SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses critical limitations in existing emotion recognition datasets—such as limited scale, imbalanced emotion distributions, misaligned modalities, and loose data partitioning—that hinder cross-dataset generalization and performance on minority emotions. To overcome these challenges, the authors introduce a large-scale multimodal dataset comprising 30,000 high-quality conversational segments, evenly distributed across seven emotion categories and integrating text, audio, and visual modalities. The dataset employs cinematic-grade strict partitioning to prevent content leakage and leverages a hybrid annotation pipeline combining pretrained models with human verification, alongside precise multimodal synchronization. This resource substantially enhances model robustness in in-domain evaluation, cross-dataset transfer, low-resource learning, and modality-transfer scenarios, demonstrating particularly significant improvements over existing methods in recognizing underrepresented emotions such as fear and disgust.
📝 Abstract
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.
Problem

Research questions and friction points this paper is trying to address.

multimodal emotion recognition
dataset bias
class imbalance
cross-dataset generalization
affective computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal emotion recognition
balanced dataset
conversation-level emotion
cross-dataset generalization
modality alignment