emotion recognition

Designs, builds, or analyzes systems that detect and characterize human emotional states from speech, facial-expression and other multimodal signals; this includes models that produce per-frame or continuous emotion trajectories, context-aware and chunk/state-based temporal representations (including emotion-state chaining), fine-grained and zero-shot emotion classification, emotion embeddings, compact on-device inference pipelines, and the evaluation/benchmarking protocols and metrics (e.g., per-class accuracy, macro-F1) used to compare methods.

emotionrecognition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current video-based affective analysis models face bottlenecks in fine-grained multimodal co-modeling of micro-expressions and speech, compounded by a scarcity of high-quality, multimodal emotional datasets. To address this, we propose the first multimodal affective analysis framework explicitly designed for micro-expression–speech alignment. Our approach introduces a two-tier annotated dataset (24K coarse-grained + 3.5K fine-grained samples), incorporating both self-supervised and human-refined annotations. We design a facial micro-expression encoder and a temporal audio modeling module, enabling cross-modal fine-grained alignment within a unified representation space. Additionally, we incorporate instruction-tuned joint optimization to simultaneously enhance emotion recognition and reasoning capabilities. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, with significant improvements in discriminative accuracy and interpretability for subtle emotions—particularly contempt and confusion.

Facial Micro-ExpressionsMultimodal Emotion AnalysisVoice Analysis

This study addresses the challenge of real-time facial emotion recognition in videos, where large inter-individual variability and subtle, continuous expression dynamics hinder performance. To tackle this, the authors propose a deep neural network that integrates multi-scale feature extraction with supervised contrastive learning. By explicitly modeling the temporal evolution of facial expressions, the method effectively captures fine-grained emotional distinctions and substantially enhances model generalization. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach achieves state-of-the-art recognition accuracy while maintaining real-time inference speed, offering robust affective perception capabilities for practical applications such as psychological counseling.

continuous emotional statesfacial expression dynamicsindividual variation in facial expressions

Human Feedback Driven Dynamic Speech Emotion Recognition

Aug 18, 2025
IF
Ilya Fedorov
🏛️ NVIDIA

Speech emotion is inherently time-varying, yet conventional methods often assume a static, single-label emotion per utterance—limiting their applicability to natural, real-time affective animation of 3D virtual agents. To address this, we propose a multi-stage training and human-feedback-driven optimization framework for dynamic speech emotion recognition. First, we model emotional mixtures using a Dirichlet distribution, enabling fine-grained, continuous emotion sequence prediction. Second, we integrate human feedback into a reinforcement learning loop to iteratively refine the model while substantially reducing reliance on dense manual annotations. Experiments demonstrate that Dirichlet-based modeling significantly outperforms sliding-window baselines, achieving a 4.2% F1-score gain on RAVDESS and comparable datasets. Moreover, annotation efficiency improves by approximately 60%. This work establishes a novel, interpretable, and optimization-friendly paradigm for dynamic emotion modeling, directly supporting expressive, real-time emotional animation in virtual humans.

Dynamic speech emotion recognition with sequential emotional labelsImproving emotion recognition through human feedback integrationModeling emotional mixtures using Dirichlet distribution approach

This study investigates the feasibility and performance of large language models (LLMs) for zero-shot emotion recognition in real-life videos. Addressing the FERV39k-DailyLife dataset, it performs both fine-grained (seven classes: Angry, Disgust, Fear, Happy, Neutral, Sad, Surprise) and coarse-grained (three classes: Negative, Neutral, Positive) emotion classification on keyframes. To enable cross-modal emotion understanding without fine-tuning or labeled data, we propose a novel multi-frame fusion prompting strategy. Experimental results show that GPT-4o-mini achieves 50% average precision on the seven-class task and 64% on the three-class task. Multi-frame fusion significantly improves robustness and reduces annotation overhead. This work departs from conventional supervised paradigms, demonstrating that LLMs—when guided by carefully designed visual prompting—can effectively infer emotional states from uncurated, naturalistic video content. It establishes a new pathway for automated, low-resource emotion analysis in real-world settings.

Large language models for emotion annotationMulti-frame approach to enhance accuracyZero-shot labeling in daily scenarios

Context-aware Deep Learning for Multi-modal Depression Detection

May 01, 2019
GL
Genevieve Lam
🏛️ Nanyang Technological University | Institute of Infocomm Research | UBTech

This study addresses automatic depression detection in clinical interview settings through a context-aware multimodal (audio + text) fusion approach. Methodologically, it introduces: (i) a novel topic-modeling–based text data augmentation strategy leveraging BERTopic and LDA; (ii) deep 1D convolutional neural networks for acoustic feature modeling and Transformer architectures for semantic textual representation; and (iii) a cross-modal alignment and adaptive fusion mechanism. Experimental results demonstrate state-of-the-art (SOTA) performance for both unimodal modalities—audio (+3.2% accuracy gain) and text—while the multimodal system achieves performance on par with the best contemporary systems. The proposed framework offers an interpretable, robust paradigm for clinical speech-text analysis under low-resource and high-noise conditions, advancing practical applicability in real-world mental health assessment.

Depression DetectionSpeech AnalysisText Analysis

Latest Papers

What's happening recently
View more

This work addresses four challenging tasks in real-world scenarios: facial expression recognition, valence-arousal estimation, action unit detection, and fine-grained violent behavior classification. The authors propose an efficient two-stage prediction framework that leverages EfficientNet-based pretrained models to extract facial embeddings, followed by confidence-thresholded frame-level predictions using multilayer perceptrons. Temporal consistency is enhanced through a sliding-window smoothing strategy. For violent behavior detection, the study systematically evaluates various pretrained architectures and video-level embedding aggregation methods. The proposed approach significantly outperforms existing baselines across all four tasks in the ABAW-10 challenge, achieving substantial gains in robustness for affective and behavioral understanding under complex conditions while maintaining high inference efficiency.

Action Unit DetectionAffective Behavior AnalysisFacial Expression Recognition

This work addresses the challenge of efficiently executing multimodal perception tasks on low-power edge devices, where existing intelligent surveillance systems struggle with both computational efficiency and context-aware resource management. We propose a real-time multimodal vision framework tailored for the Raspberry Pi 5, integrating YOLOv8n for object detection, a customized FaceNet module for face recognition, and DeepFace for emotion classification. A context-triggered adaptive runtime scheduler dynamically activates subtasks only when needed, enabling effective task coordination while substantially reducing computational load. Experimental results demonstrate a 65% reduction in computational overhead, with an object detection AP of 0.861, 88% face recognition accuracy, and an emotion classification AUC up to 0.97, achieving an overall inference speed of 5.6 FPS. These findings validate the feasibility of deploying complex multimodal AI pipelines efficiently on cost-constrained edge hardware.

adaptive schedulingedge computinglow-power

This work addresses the limitation of existing affective understanding benchmarks, which typically treat emotions as static states and thus fail to evaluate multimodal large language models’ capacity to model dynamic emotional evolution and state transitions within social contexts. To bridge this gap, the authors propose EmoTrans—the first multitask evaluation benchmark focused on emotional dynamics—comprising 12 real-world scenarios, 1,000 annotated videos, and over 3,000 structured question-answer pairs. EmoTrans encompasses four progressively challenging tasks: emotion change detection, state recognition, transition reasoning, and next-emotion prediction. Systematic evaluation of 18 state-of-the-art models reveals that while current approaches perform reasonably well on coarse-grained change detection, they exhibit significant deficiencies in fine-grained dynamic modeling and complex social scenarios, with reasoning-enhancement strategies yielding inconsistent performance gains.

benchmarkemotion dynamicsemotion transition

Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies

Dec 23, 2025
JH
Jing Han
🏛️ University of Cambridge | Hunan University | Imperial College London | TUM University Hospital | Munich Center for Machine Learning | Munich Data Science Institute

Prior work lacks systematic evaluation of multimodal large language models (MLLMs) on fine-grained emotion understanding in open-vocabulary multimodal emotion recognition (MER-OV). Method: We introduce the first large-scale MER-OV benchmark—built upon the OV-MERD dataset—and comprehensively evaluate 19 state-of-the-art MLLMs across audio, video, and text modalities. We propose a multidimensional evaluation paradigm covering reasoning analysis, modality fusion, context utilization, and prompt engineering. Contribution/Results: Our study reveals that two-stage trimodal fusion is optimal, with video contributing most to performance; open- and closed-source MLLMs exhibit negligible performance gaps. Our framework achieves new state-of-the-art results on MER-OV. We publicly release code, models, and practical guidelines to advance interpretable, fine-grained affective AI.

Benchmarking multimodal large models for open-vocabulary emotion recognitionEvaluating fine-grained emotion understanding across diverse model architecturesIdentifying optimal fusion strategies for audio, video, and text modalities

This work addresses emotion recognition in conversational scenarios by effectively integrating multimodal information to enhance performance. We propose a lightweight multimodal baseline system that combines a Transformer-based text classifier with a self-supervised speech representation model, employing a simple late-fusion strategy for emotion prediction. Experimental results on the SemEval-2024 Task 3 dataset demonstrate that, under constrained training conditions, our multimodal approach significantly outperforms unimodal models. By providing a transparent and reproducible benchmark system, this study establishes a reliable foundation for future research in multimodal emotion recognition within dialogue contexts.

baselineconversational emotionemotion recognition

Hot Scholars

HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
ZZ

Zhixian Zhao

Northwestern Polytechnical University
Emotion Speech RecognitionUnderstanding and Generation
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
BS

Berrak Sisman

Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice Conversion
ZL

Zheng Lian

Associate Professor, IEEE/CCF Senior Member, Institute of Automation, Chinese Academy of Sciences
Affective ComputingSentiment AnalysisMachine Learning