facial emotion recognition

Designs, builds, or evaluates systems that detect and classify human emotional states from facial appearance and expressions in images or video. Work covers feature extraction (landmarks, appearance, motion), model training and validation for expression/emotion classification or intensity estimation, annotation and label-schema management, and analysis of robustness to pose, lighting, occlusion, and temporal dynamics.

facialemotionrecognition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Facial Landmark Visualization and Emotion Recognition Through Neural Networks

Jun 20, 2025
IJ
Israel Ju'arez-Jim'enez
🏛️ Unidad Profesional Interdisciplinaria de Ingenier'ia Campus Tlaxcala IPN

This study addresses facial emotion recognition in human-computer interaction, focusing on data quality assessment and feature representation optimization. To this end, we propose a novel boxplot-based visualization method leveraging facial landmarks for outlier detection in facial datasets—the first such application. We systematically compare absolute-coordinate landmarks against neutral-to-peak displacement features, providing the first empirical evidence that displacement features significantly outperform absolute coordinates in emotion classification. Landmarks are extracted using dlib and MMPose; classification is performed via CNN and Random Forest models. Results demonstrate that CNN substantially surpasses Random Forest in accuracy; moreover, displacement features markedly enhance model robustness and cross-dataset generalization. This work contributes an interpretable, landmark-driven data quality control tool and establishes a superior feature paradigm—displacement-based representation—for facial emotion recognition.

Compare absolute vs displacement facial landmark featuresDevelop facial landmark visualization for dataset outlier detectionImprove emotion recognition using neural networks over random forests

Current facial emotion recognition systems exhibit significant limitations in cross-dataset generalization, handling class imbalance, and distinguishing subtle emotional states—such as disgust versus anger. This study presents the first systematic evaluation of three state-of-the-art deep learning models across three large-scale, diverse datasets, revealing a substantial performance drop when models are tested on unseen data. The work further investigates how inter-dataset label inconsistencies and variations in annotation difficulty critically undermine model robustness. By identifying these key weaknesses in existing approaches, the study provides empirical evidence and clear directions for future research aimed at improving cross-domain generalization and fine-grained emotion recognition.

Cross-dataset GeneralizationDataset DiscrepancyEmotion Differentiation

Faces of the Mind: Unveiling Mental Health States Through Facial Expressions in 11,427 Adolescents

May 30, 2024
XX
Xiao Xu
🏛️ Nanjing Medical University | The Affiliated Brain Hospital of Nanjing Medical University

Existing depression and anxiety recognition models suffer from limited generalizability in real-world settings due to small-sample constraints and insufficient modeling of clinically relevant behavioral biomarkers. Method: We introduce the first large-scale, multimodal dataset for adolescent mental health (n = 11,427), integrating standardized facial videos with validated psychological assessments. We systematically model novel oculomotor biomarkers—including pupillary dynamics and gaze direction—and propose a multi-granularity framework addressing symptom heterogeneity: it hybridizes tree-based and deep learning classifiers, incorporates facial action units and eye-tracking features, and uncovers emotion subtypes via clustering. Results: Pupillary fluctuations significantly correlate with perceived stress levels (p < 0.001). Our model achieves an AUC of 0.82 on real-world data—outperforming state-of-the-art methods by 12.6%. Critically, we empirically characterize the performance degradation of few-shot models under large-scale deployment for the first time, establishing a new paradigm for clinically interpretable digital phenotyping.

Assessing mental health via facial expressions in large adolescent datasetsImproving model accuracy by quantifying and managing dataset heterogeneityOvercoming symptom heterogeneity in machine learning models for mood disorders

OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis

Jun 03, 2025
JH
Jiewen Hu
🏛️ Carnegie Mellon University | Massachusetts Institute of Technology

Facial behavior analysis—encompassing facial landmark localization, action unit recognition, gaze estimation, and emotion recognition—faces significant challenges in cross-scenario generalization and computational efficiency. To address these, this work introduces the first lightweight unified multi-task model for facial behavior analysis. Our approach employs a shared backbone network coupled with task-specific lightweight heads, end-to-end joint optimization, and a cross-domain robust training paradigm to enhance generalization across diverse demographics, head poses, illumination conditions, and image resolutions. The resulting model achieves state-of-the-art or competitive accuracy on multiple benchmarks while reducing inference latency by 40% and memory footprint by 55%. All code and pre-trained models are fully open-sourced, enabling plug-and-play deployment and community-driven extension.

Develops a lightweight multitask system for facial behavior analysisEnables real-time operation without specialized hardware requirementsImproves performance in facial landmark and action unit detection

Contextual Emotion Recognition using Large Vision Language Models

May 14, 2024
YE
Yasaman Etesam
🏛️ Simon Fraser University

Existing facial-expression-only approaches to human emotion recognition in real-world scenarios suffer from poor generalization due to insufficient contextual grounding. Method: This paper proposes a contextualized emotion understanding framework that jointly models body pose, environmental context, and commonsense reasoning. It systematically evaluates large vision-language models (VLMs) for fine-grained contextual emotion recognition under zero-shot and few-shot fine-tuning settings, introducing a dual-path multimodal architecture: (i) end-to-end VLM-based joint reasoning and (ii) a two-stage pipeline comprising image captioning followed by pure language-model inference. Contribution/Results: On the EMOTIC benchmark, fine-tuning only a small-scale VLM surpasses state-of-the-art unimodal and multimodal baselines. The method establishes a novel paradigm for embodied agents to achieve robust, context-sensitive affective perception and interaction.

Artificial IntelligenceComputer VisionEmotion Recognition

Latest Papers

What's happening recently
View more

This study addresses the challenge of real-time facial emotion recognition in videos, where large inter-individual variability and subtle, continuous expression dynamics hinder performance. To tackle this, the authors propose a deep neural network that integrates multi-scale feature extraction with supervised contrastive learning. By explicitly modeling the temporal evolution of facial expressions, the method effectively captures fine-grained emotional distinctions and substantially enhances model generalization. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach achieves state-of-the-art recognition accuracy while maintaining real-time inference speed, offering robust affective perception capabilities for practical applications such as psychological counseling.

continuous emotional statesfacial expression dynamicsindividual variation in facial expressions

This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.

Address semantic correctness and system constraints in video-to-artifact pipelinesDetect visible people and emotions from video frames using vision-language modelsGenerate structured bounding box outputs with prompt-conditioned attributes

This work addresses four challenging tasks in real-world scenarios: facial expression recognition, valence-arousal estimation, action unit detection, and fine-grained violent behavior classification. The authors propose an efficient two-stage prediction framework that leverages EfficientNet-based pretrained models to extract facial embeddings, followed by confidence-thresholded frame-level predictions using multilayer perceptrons. Temporal consistency is enhanced through a sliding-window smoothing strategy. For violent behavior detection, the study systematically evaluates various pretrained architectures and video-level embedding aggregation methods. The proposed approach significantly outperforms existing baselines across all four tasks in the ABAW-10 challenge, achieving substantial gains in robustness for affective and behavioral understanding under complex conditions while maintaining high inference efficiency.

Action Unit DetectionAffective Behavior AnalysisFacial Expression Recognition

Existing facial expression recognition datasets predominantly rely on static images, basic emotion categories, or single-label annotations, limiting their ability to capture the dynamics of facial expressions and the diversity of human perception. To address this, this work proposes Chehre—a novel, anonymized facial expression video dataset that leverages emojis as both expressive prompts and annotation primitives. The dataset employs facial motion transfer to generate synthetic videos for privacy preservation and collects multi-label emotion distributions via crowdsourcing. The study formulates a distributional facial expression recognition task and evaluates vision-language models under character-specific emoji prompts through multi-prediction assessment. Experiments reveal that state-of-the-art models achieve only 32.5% Top-1 accuracy on dominant emotion recognition, and their Spread Ratio on the distributional task remains substantially below human performance, highlighting the challenge and research potential of this new benchmark.

distributional annotationdynamic expressionsemoji-prompted

This study addresses the efficient assessment of consumer acceptance of branded products in supermarket or hypermarket settings by leveraging real-time analysis of facial expressions during product selection. To this end, an enhanced Harris corner detection algorithm is proposed, which significantly reduces computational time complexity while preserving high accuracy in facial expression recognition. The method optimizes the extraction of facial feature points, thereby improving the overall efficiency of the expression recognition pipeline and enhancing its suitability for real-time deployment in authentic retail environments. Experimental results demonstrate that the proposed algorithm outperforms existing approaches in corner detection speed, achieving a favorable balance between accuracy and real-time performance, thus effectively supporting public acceptance evaluation of products based on spontaneous facial expressions.

Brand EvaluationFacial Expression DetectionProduct Review

Hot Scholars

AS

Ali Safa

Assistant Professor, HBKU
Neuro-RoboticsHuman-Robot InteractionSocial RoboticsAI
FM

Farida Mohsen

Postdoctoral Researcher, HBKU, College of Science and Engineering
Artificial IntelligenceMedical AImedical imagingNLP
ZL

Zhengyu Li

Peking University
Quantum Cryptography
AM

Alex Mihailidis

Professor, Department of Occupational Science & Occupational Therapy, University of Toronto
artificial intelligencepervasive computingolder adultsassistive technology