Score
Designs, builds, or analyzes systems that detect and characterize human emotional states from speech, facial-expression and other multimodal signals; this includes models that produce per-frame or continuous emotion trajectories, context-aware and chunk/state-based temporal representations (including emotion-state chaining), fine-grained and zero-shot emotion classification, emotion embeddings, compact on-device inference pipelines, and the evaluation/benchmarking protocols and metrics (e.g., per-class accuracy, macro-F1) used to compare methods.
Current video-based affective analysis models face bottlenecks in fine-grained multimodal co-modeling of micro-expressions and speech, compounded by a scarcity of high-quality, multimodal emotional datasets. To address this, we propose the first multimodal affective analysis framework explicitly designed for micro-expression–speech alignment. Our approach introduces a two-tier annotated dataset (24K coarse-grained + 3.5K fine-grained samples), incorporating both self-supervised and human-refined annotations. We design a facial micro-expression encoder and a temporal audio modeling module, enabling cross-modal fine-grained alignment within a unified representation space. Additionally, we incorporate instruction-tuned joint optimization to simultaneously enhance emotion recognition and reasoning capabilities. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, with significant improvements in discriminative accuracy and interpretability for subtle emotions—particularly contempt and confusion.
This study addresses the challenge of real-time facial emotion recognition in videos, where large inter-individual variability and subtle, continuous expression dynamics hinder performance. To tackle this, the authors propose a deep neural network that integrates multi-scale feature extraction with supervised contrastive learning. By explicitly modeling the temporal evolution of facial expressions, the method effectively captures fine-grained emotional distinctions and substantially enhances model generalization. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach achieves state-of-the-art recognition accuracy while maintaining real-time inference speed, offering robust affective perception capabilities for practical applications such as psychological counseling.
Speech emotion is inherently time-varying, yet conventional methods often assume a static, single-label emotion per utterance—limiting their applicability to natural, real-time affective animation of 3D virtual agents. To address this, we propose a multi-stage training and human-feedback-driven optimization framework for dynamic speech emotion recognition. First, we model emotional mixtures using a Dirichlet distribution, enabling fine-grained, continuous emotion sequence prediction. Second, we integrate human feedback into a reinforcement learning loop to iteratively refine the model while substantially reducing reliance on dense manual annotations. Experiments demonstrate that Dirichlet-based modeling significantly outperforms sliding-window baselines, achieving a 4.2% F1-score gain on RAVDESS and comparable datasets. Moreover, annotation efficiency improves by approximately 60%. This work establishes a novel, interpretable, and optimization-friendly paradigm for dynamic emotion modeling, directly supporting expressive, real-time emotional animation in virtual humans.
This study investigates the feasibility and performance of large language models (LLMs) for zero-shot emotion recognition in real-life videos. Addressing the FERV39k-DailyLife dataset, it performs both fine-grained (seven classes: Angry, Disgust, Fear, Happy, Neutral, Sad, Surprise) and coarse-grained (three classes: Negative, Neutral, Positive) emotion classification on keyframes. To enable cross-modal emotion understanding without fine-tuning or labeled data, we propose a novel multi-frame fusion prompting strategy. Experimental results show that GPT-4o-mini achieves 50% average precision on the seven-class task and 64% on the three-class task. Multi-frame fusion significantly improves robustness and reduces annotation overhead. This work departs from conventional supervised paradigms, demonstrating that LLMs—when guided by carefully designed visual prompting—can effectively infer emotional states from uncurated, naturalistic video content. It establishes a new pathway for automated, low-resource emotion analysis in real-world settings.
This study addresses automatic depression detection in clinical interview settings through a context-aware multimodal (audio + text) fusion approach. Methodologically, it introduces: (i) a novel topic-modeling–based text data augmentation strategy leveraging BERTopic and LDA; (ii) deep 1D convolutional neural networks for acoustic feature modeling and Transformer architectures for semantic textual representation; and (iii) a cross-modal alignment and adaptive fusion mechanism. Experimental results demonstrate state-of-the-art (SOTA) performance for both unimodal modalities—audio (+3.2% accuracy gain) and text—while the multimodal system achieves performance on par with the best contemporary systems. The proposed framework offers an interpretable, robust paradigm for clinical speech-text analysis under low-resource and high-noise conditions, advancing practical applicability in real-world mental health assessment.
This work addresses four challenging tasks in real-world scenarios: facial expression recognition, valence-arousal estimation, action unit detection, and fine-grained violent behavior classification. The authors propose an efficient two-stage prediction framework that leverages EfficientNet-based pretrained models to extract facial embeddings, followed by confidence-thresholded frame-level predictions using multilayer perceptrons. Temporal consistency is enhanced through a sliding-window smoothing strategy. For violent behavior detection, the study systematically evaluates various pretrained architectures and video-level embedding aggregation methods. The proposed approach significantly outperforms existing baselines across all four tasks in the ABAW-10 challenge, achieving substantial gains in robustness for affective and behavioral understanding under complex conditions while maintaining high inference efficiency.
This work addresses the challenge of efficiently executing multimodal perception tasks on low-power edge devices, where existing intelligent surveillance systems struggle with both computational efficiency and context-aware resource management. We propose a real-time multimodal vision framework tailored for the Raspberry Pi 5, integrating YOLOv8n for object detection, a customized FaceNet module for face recognition, and DeepFace for emotion classification. A context-triggered adaptive runtime scheduler dynamically activates subtasks only when needed, enabling effective task coordination while substantially reducing computational load. Experimental results demonstrate a 65% reduction in computational overhead, with an object detection AP of 0.861, 88% face recognition accuracy, and an emotion classification AUC up to 0.97, achieving an overall inference speed of 5.6 FPS. These findings validate the feasibility of deploying complex multimodal AI pipelines efficiently on cost-constrained edge hardware.
This work addresses the limitation of existing affective understanding benchmarks, which typically treat emotions as static states and thus fail to evaluate multimodal large language models’ capacity to model dynamic emotional evolution and state transitions within social contexts. To bridge this gap, the authors propose EmoTrans—the first multitask evaluation benchmark focused on emotional dynamics—comprising 12 real-world scenarios, 1,000 annotated videos, and over 3,000 structured question-answer pairs. EmoTrans encompasses four progressively challenging tasks: emotion change detection, state recognition, transition reasoning, and next-emotion prediction. Systematic evaluation of 18 state-of-the-art models reveals that while current approaches perform reasonably well on coarse-grained change detection, they exhibit significant deficiencies in fine-grained dynamic modeling and complex social scenarios, with reasoning-enhancement strategies yielding inconsistent performance gains.
Prior work lacks systematic evaluation of multimodal large language models (MLLMs) on fine-grained emotion understanding in open-vocabulary multimodal emotion recognition (MER-OV). Method: We introduce the first large-scale MER-OV benchmark—built upon the OV-MERD dataset—and comprehensively evaluate 19 state-of-the-art MLLMs across audio, video, and text modalities. We propose a multidimensional evaluation paradigm covering reasoning analysis, modality fusion, context utilization, and prompt engineering. Contribution/Results: Our study reveals that two-stage trimodal fusion is optimal, with video contributing most to performance; open- and closed-source MLLMs exhibit negligible performance gaps. Our framework achieves new state-of-the-art results on MER-OV. We publicly release code, models, and practical guidelines to advance interpretable, fine-grained affective AI.
This work addresses emotion recognition in conversational scenarios by effectively integrating multimodal information to enhance performance. We propose a lightweight multimodal baseline system that combines a Transformer-based text classifier with a self-supervised speech representation model, employing a simple late-fusion strategy for emotion prediction. Experimental results on the SemEval-2024 Task 3 dataset demonstrate that, under constrained training conditions, our multimodal approach significantly outperforms unimodal models. By providing a transparent and reproducible benchmark system, this study establishes a reliable foundation for future research in multimodal emotion recognition within dialogue contexts.