emotion detection

Detecting and modeling expressed emotions or sentiment from modalities such as facial expressions or text—potentially as continuous signals—and relating those measurements to outcomes like engagement, coping styles, or anthropomorphization.

emotiondetection

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Measuring Non-Typical Emotions for Mental Health: A Survey of Computational Approaches

Mar 09, 2024
PK
Puneet Kumar
🏛️ University of Oulu | Zhejiang University

This paper addresses the challenges of recognizing three complex psychological states—stress, depression, and engagement—and the lack of a unified modeling framework for their joint analysis. Methodologically, it conducts the first systematic review simultaneously covering all three states, integrating multimodal data (speech, text, physiological signals), feature- and decision-level fusion strategies, and both machine learning and deep learning models; it performs cross-benchmark comparative analysis of state-of-the-art performance on major datasets including DAIC-WOZ, AVEC, and RECOLA. Key contributions include: (1) proposing the first unified taxonomy and technology evolution timeline for these three states; (2) establishing a general computational analysis pipeline; and (3) identifying critical bottlenecks in model interpretability, cross-population generalizability, and privacy preservation. The work provides both theoretical foundations and practical guidelines for computational modeling of atypical psychological states.

Analyzing stress, depression, and engagement computationally is complex and less common.Exploring the interplay between stress, depression, and their impact on daily engagement.Surveying computational methods, datasets, and applications for mental health analysis.

Empathy Detection from Text, Audiovisual, Audio or Physiological Signals: A Systematic Review of Task Formulations and Machine Learning Methods

Oct 30, 2023
MR
Md Rakibul Hasan
🏛️ Curtin University | BRAC University | The Australian National University | University of ´OBuda

Empathic response detection faces challenges including ill-defined task formulations, fragmented multimodal modeling, and the absence of systematic surveys. This paper conducts a cross-modal systematic review of 62 high-quality studies. Methodologically, it unifies network architecture design principles along modality dimensions for the first time, proposes a standardized three-tier interaction framework (individual, dyadic, and group) and a five-category task taxonomy—including local/global empathy recognition and emotional contagion detection. It constructs the first structured empathy detection knowledge graph, integrating four input modalities (text, audio, audiovisual, and physiological signals), twelve publicly available datasets, and seven reproducible codebases. By synergizing NLP, audiovisual modeling, speech emotion analysis, and time-frequency physiological signal processing—augmented with cross-modal contrastive learning and meta-analysis—the study identifies critical research gaps, establishing both theoretical foundations and practical guidelines for robust empathic computing.

Analyzing task formulations and ML approaches in empathy detectionIdentifying challenges and gaps in Affective Computing-based empathy researchSystematically reviewing empathy detection methods from multiple modalities

Must-Read Papers

Most classic and influential ideas
View more

Existing research often examines textual, visual, or audio modalities of short videos in isolation, failing to uncover how their interplay influences user engagement. This work proposes the first reproducible and interpretable multimodal analysis framework that integrates automated feature extraction with Shapley-value-based attribution to systematically investigate how multimodal interactions affect view counts in TikTok content related to social anxiety disorder. The study reveals that facial expressions are more predictive than textual sentiment, that informational content garners greater attention than emotional support, and that multimodal synergies exhibit strong threshold-dependent effects—thereby transcending the limitations of conventional unimodal analyses.

cross-modal interactionmental health discoursemultimodal communication

Bridging Discrete and Continuous: A Multimodal Strategy for Complex Emotion Detection

Sep 12, 2024
JJ
Jiehui Jia
🏛️ Queen Mary University of London

To address the limitations of conventional emotion recognition in human-computer interaction—namely, oversimplified modeling and insufficient coverage of discrete emotion categories—this paper proposes a multimodal continuous emotion modeling framework. It fuses facial expression, prosodic, and textual transcription features to construct a continuous emotion representation in the three-dimensional Valence-Arousal-Dominance (VAD) space. A novel contribution is the adaptive mapping of discrete emotion labels to the VAD space via K-means clustering, enabling open-vocabulary emotion generation. Additionally, a cross-modal feature alignment mechanism and a joint classifier are designed to enhance multimodal fusion. Evaluated on the MER2024 Chinese film-and-television dataset, the method achieves a 42% improvement in emotion vocabulary coverage and an accuracy of 91.3%, effectively balancing emotional diversity and discriminative precision.

Bridging discrete and continuous emotion recognition models effectivelyDetecting complex emotions via multimodal facial, vocal, and text dataMapping emotions in a continuous Valence-Arousal-Dominance (VAD) space

Exploring Machine Learning and Language Models for Multimodal Depression Detection

Aug 28, 2025
JS
Javier Si Zhao Hong
🏛️ Singapore Institute of Technology | Duke Kunshan University

This study addresses automated multimodal depression detection by proposing a cross-modal modeling framework integrating audio, visual, and textual features. Methodologically, it introduces a synergistic architecture combining XGBoost for shallow temporal statistical feature learning, Transformer networks for intra-modal long-range dependency modeling, and large language models (LLMs) for enhanced textual semantic and contextual understanding, augmented by modality alignment and weighted fusion strategies. Comprehensive evaluation on public multimodal depression datasets demonstrates that the proposed fusion model significantly outperforms unimodal baselines—achieving AUC improvements of 4.2–7.8 percentage points—thereby validating the complementary strengths of heterogeneous models and the efficacy of multimodal representation learning. The work establishes a novel, interpretable, and robust paradigm for mental health status assessment, offering a reproducible technical pathway grounded in principled multimodal integration.

Comparing XGBoost, transformers and LLMs across modalitiesDetecting depression using multimodal machine learning modelsEvaluating model performance on depression-related signal detection

Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values

Feb 08, 2025
VM
Vaibhav Mehra
🏛️ Universidade Lusófona | University of Cambridge

This work investigates whether large language models (LLMs) can perform affective semantic understanding solely from valence-arousal (VA) numerical representations—without visual input. We evaluate zero-shot emotion classification and prompt-driven semantic description generation on the IIMI (basic emotions) and Emotic (complex emotions) datasets; VA features are extracted by FaceChannel, and models including GPT and Llama are assessed. Results show that LLMs exhibit limited performance on VA-based classification—particularly in distinguishing non-binary emotions—yet achieve strong performance in semantic description generation (BLEU-4 = 0.62; human-rated semantic similarity = 4.3/5), demonstrating robust mapping from abstract affective dimensions to natural language. To our knowledge, this is the first study to empirically validate that LLMs can directly interpret non-visual, structured affective representations. The findings establish a novel paradigm for lightweight, bias-mitigated affective computing grounded in interpretable, modality-agnostic emotional features.

LLMs classify facial expressions into basic and complex emotions.LLMs generate semantic descriptions of facial expressions.LLMs infer facial expressions from Valence-Arousal values.

Annotation and modeling of emotions in a textual corpus: an evaluative approach

Sep 01, 2025
JN
Jonas Noblet
🏛️ Université Grenoble Alpes | Société Ixiade

This study addresses the substantial inter-annotator disagreement and low stability observed in sentiment annotation. Grounded in Appraisal Theory, it systematically investigates the textual features of sentiment expression and their computability. Using a manually annotated industrial corpus, we construct a fine-grained sentiment dataset and train language models to emulate human annotation behavior. Methodologically, we move beyond mainstream sentiment classification paradigms by incorporating appraisal-oriented semantic structures—namely, Attitude, Engagement, and Graduation—which remain underexplored in computational linguistics. This enables us to uncover stable statistical patterns and language-driven mechanisms underlying annotation discrepancies. Experiments demonstrate that our model effectively discriminates among distinct appraisal-based sentiment contexts, achieving significant improvements over baselines in cross-context generalization and fine-grained linguistic cue modeling. Our work establishes a theory-informed modeling framework for sentiment computation and provides an interpretable, appraisal-grounded evaluation benchmark.

Demonstrating statistical stability in subjective emotional annotationsEvaluating language models' capability to distinguish emotional situationsModeling emotion annotation disagreements in industrial text corpus

Latest Papers

What's happening recently
View more

This study addresses key challenges in multimodal emotion recognition—namely, the scarcity of affectively annotated data, semantic gaps across modalities, and opaque reasoning processes—by proposing a novel paradigm termed “MER-with-LLMs.” The work systematically constructs a large language model (LLM)-based framework that integrates three core components: affective data augmentation, cross-modal alignment, and interpretable reasoning. It establishes, for the first time, a comprehensive taxonomy and research roadmap for leveraging LLMs in general-purpose affective intelligence. Through a thorough survey of existing approaches, the paper delineates a clear academic landscape, elucidating the field’s developmental trajectory, critical bottlenecks, and emerging directions, thereby advancing multimodal emotion recognition toward greater structural coherence and interpretability.

Affective GapEmotion AnnotationInterpretability

Fluent but Unfeeling: The Emotional Blind Spots of Language Models

Sep 11, 2025
BS
Bangzhao Shu
🏛️ Northeastern University | UC San Diego | University of Massachusetts Amherst | Independent Researcher

This study investigates the alignment between large language models (LLMs) and human self-reported emotions in fine-grained sentiment recognition—moving beyond conventional coarse-grained classification. To this end, we introduce EXPRESS, the first benchmark dataset for fine-grained self-reported emotion analysis, comprising 251 authentic user-emotion labels. Grounded in classical emotion theory, we decompose model predictions into eight basic emotion categories, establishing an interpretable, fine-grained alignment evaluation framework. Through systematic experiments across diverse prompting strategies, we evaluate leading LLMs and find that while they generate theoretically plausible emotion terms, their capacity to capture context-dependent, nuanced affective states remains substantially inferior to human performance. Our core contributions are threefold: (1) the EXPRESS dataset; (2) a novel fine-grained emotion decomposition and alignment evaluation paradigm; and (3) empirical characterization of the fundamental limits of LLMs’ emotional alignment capability.

Assessing nuanced emotional expressions beyond predefined categoriesEvaluating LLMs' fine-grained emotion alignment with humansTesting contextual emotion cue capture in language models

Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing

Jun 08, 2025
YT
Yuanhe Tian
🏛️ University of Washington | Sichuan University | People's Daily Online | University of Science and Technology of China

Existing methods struggle to jointly model the similarity and conflict between affective signals across modalities (e.g., image and text). To address this, we propose the first large language model (LLM)-based representation decomposition framework that explicitly disentangles cross-modal shared affective components from modality-specific cues. Our approach employs an attention-driven dynamic soft prompting mechanism for adaptive multimodal fusion and guides the fine-tuning of multimodal LLMs. The architecture integrates a pretrained multimodal encoder, a representation decomposition network, and a soft prompt generation module. Evaluated on three benchmark tasks—sentiment analysis, emotion recognition, and hate meme detection—our method consistently outperforms state-of-the-art approaches, achieving average accuracy gains of 2.3–4.1%. Notably, it is the first to jointly model both congruent and adversarial evidence in multimodal affective understanding.

Decomposing multi-modal representations into shared and unique componentsEnhancing affective computing via dynamic LLM prompts for cross-modal fusionImproving emotion recognition from conflicting visual and textual cues

Multimodal affective computing (MAC) suffers from unstable performance and insufficient understanding of how model architectures and data characteristics jointly influence performance. To address this, we propose a hybrid optimization framework integrating generative knowledge prompting, cross-modal alignment, and supervised fine-tuning, accompanied by a systematic benchmark for comprehensive evaluation of state-of-the-art open-source multimodal large language models (MLLMs) on audio-visual-text fusion-based emotion recognition. Extensive experiments across multiple standard benchmarks demonstrate significant improvements in end-to-end emotion analysis accuracy and robustness. This work provides the first empirical characterization of the synergistic interplay between architectural design choices and data properties in multimodal emotion understanding, establishing an interpretable and reproducible paradigm for MAC model development. The implementation is publicly available.

Analyze impact of model architectures on affective analysisEnhance MLLMs' emotion recognition via generative knowledge promptingEvaluate MLLMs' performance on multimodal affective computing tasks

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language

May 20, 2025
PC
Phoebe Chua
🏛️ MIT | National University of Singapore | The University of Tokyo | Gallaudet University

Existing sign language emotion recognition research is severely limited due to the dual functional role of facial expressions and manual signs—both encode grammatical structure and emotional content—leading to feature coupling and annotation ambiguity. Method: We introduce the first fine-grained annotated American Sign Language (ASL) emotion video dataset, comprising 200 multimodal samples, with collaborative labeling by three Deaf interpreters across six discrete emotions, sentiment polarity, and open-ended emotional cue descriptions. We systematically disentangle syntax- and emotion-shared facial and manual features, establishing the first ASL emotion recognition benchmark. Contribution/Results: We propose a ViT+LSTM+CLIP multimodal fusion baseline, achieving 58.3% accuracy on emotion classification—substantially outperforming random chance. The dataset is publicly released on Hugging Face, addressing a critical gap in sign language affective computing.

Differentiating emotional and grammatical functions in sign languageLack of labeled datasets for sign language emotion recognitionUnderstanding emotional indicators in American Sign Language

Hot Scholars

ZL

Zheng Lian

Associate Professor, IEEE/CCF Senior Member, Institute of Automation, Chinese Academy of Sciences
Affective ComputingSentiment AnalysisMachine Learning
XP

Xiaojiang Peng

Shenzhen Technology University
Computer VisionFacial Expression RecognitionMultimodal Emotion Recognition
HG

Hatice Gunes

Full Professor of Affective Intelligence & Robotics, University of Cambridge
Artificial IntelligenceAffective AIHealth AIAI Fairness
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
GZ

Guoying Zhao

Academy Professor, IEEE Fellow, Professor of Computer Science and Engineering, University of Oulu
Affective ComputingArtificial IntelligenceComputer VisionPattern Recognition