Score
Designs, builds, or evaluates systems that detect and classify human emotional states from facial appearance and expressions in images or video. Work covers feature extraction (landmarks, appearance, motion), model training and validation for expression/emotion classification or intensity estimation, annotation and label-schema management, and analysis of robustness to pose, lighting, occlusion, and temporal dynamics.
本文综述了多模态面部状态分析,通过整合视觉、音频等信息解决传统单模态方法的环境敏感性和弱解释性问题,利用多任务学习提升表情识别精度与泛化能力。
This study addresses facial emotion recognition in human-computer interaction, focusing on data quality assessment and feature representation optimization. To this end, we propose a novel boxplot-based visualization method leveraging facial landmarks for outlier detection in facial datasets—the first such application. We systematically compare absolute-coordinate landmarks against neutral-to-peak displacement features, providing the first empirical evidence that displacement features significantly outperform absolute coordinates in emotion classification. Landmarks are extracted using dlib and MMPose; classification is performed via CNN and Random Forest models. Results demonstrate that CNN substantially surpasses Random Forest in accuracy; moreover, displacement features markedly enhance model robustness and cross-dataset generalization. This work contributes an interpretable, landmark-driven data quality control tool and establishes a superior feature paradigm—displacement-based representation—for facial emotion recognition.
Current facial emotion recognition systems exhibit significant limitations in cross-dataset generalization, handling class imbalance, and distinguishing subtle emotional states—such as disgust versus anger. This study presents the first systematic evaluation of three state-of-the-art deep learning models across three large-scale, diverse datasets, revealing a substantial performance drop when models are tested on unseen data. The work further investigates how inter-dataset label inconsistencies and variations in annotation difficulty critically undermine model robustness. By identifying these key weaknesses in existing approaches, the study provides empirical evidence and clear directions for future research aimed at improving cross-domain generalization and fine-grained emotion recognition.
Existing depression and anxiety recognition models suffer from limited generalizability in real-world settings due to small-sample constraints and insufficient modeling of clinically relevant behavioral biomarkers. Method: We introduce the first large-scale, multimodal dataset for adolescent mental health (n = 11,427), integrating standardized facial videos with validated psychological assessments. We systematically model novel oculomotor biomarkers—including pupillary dynamics and gaze direction—and propose a multi-granularity framework addressing symptom heterogeneity: it hybridizes tree-based and deep learning classifiers, incorporates facial action units and eye-tracking features, and uncovers emotion subtypes via clustering. Results: Pupillary fluctuations significantly correlate with perceived stress levels (p < 0.001). Our model achieves an AUC of 0.82 on real-world data—outperforming state-of-the-art methods by 12.6%. Critically, we empirically characterize the performance degradation of few-shot models under large-scale deployment for the first time, establishing a new paradigm for clinically interpretable digital phenotyping.
Facial behavior analysis—encompassing facial landmark localization, action unit recognition, gaze estimation, and emotion recognition—faces significant challenges in cross-scenario generalization and computational efficiency. To address these, this work introduces the first lightweight unified multi-task model for facial behavior analysis. Our approach employs a shared backbone network coupled with task-specific lightweight heads, end-to-end joint optimization, and a cross-domain robust training paradigm to enhance generalization across diverse demographics, head poses, illumination conditions, and image resolutions. The resulting model achieves state-of-the-art or competitive accuracy on multiple benchmarks while reducing inference latency by 40% and memory footprint by 55%. All code and pre-trained models are fully open-sourced, enabling plug-and-play deployment and community-driven extension.
Existing facial-expression-only approaches to human emotion recognition in real-world scenarios suffer from poor generalization due to insufficient contextual grounding. Method: This paper proposes a contextualized emotion understanding framework that jointly models body pose, environmental context, and commonsense reasoning. It systematically evaluates large vision-language models (VLMs) for fine-grained contextual emotion recognition under zero-shot and few-shot fine-tuning settings, introducing a dual-path multimodal architecture: (i) end-to-end VLM-based joint reasoning and (ii) a two-stage pipeline comprising image captioning followed by pure language-model inference. Contribution/Results: On the EMOTIC benchmark, fine-tuning only a small-scale VLM surpasses state-of-the-art unimodal and multimodal baselines. The method establishes a novel paradigm for embodied agents to achieve robust, context-sensitive affective perception and interaction.
This study addresses the challenge of real-time facial emotion recognition in videos, where large inter-individual variability and subtle, continuous expression dynamics hinder performance. To tackle this, the authors propose a deep neural network that integrates multi-scale feature extraction with supervised contrastive learning. By explicitly modeling the temporal evolution of facial expressions, the method effectively captures fine-grained emotional distinctions and substantially enhances model generalization. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach achieves state-of-the-art recognition accuracy while maintaining real-time inference speed, offering robust affective perception capabilities for practical applications such as psychological counseling.
This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.
This work addresses four challenging tasks in real-world scenarios: facial expression recognition, valence-arousal estimation, action unit detection, and fine-grained violent behavior classification. The authors propose an efficient two-stage prediction framework that leverages EfficientNet-based pretrained models to extract facial embeddings, followed by confidence-thresholded frame-level predictions using multilayer perceptrons. Temporal consistency is enhanced through a sliding-window smoothing strategy. For violent behavior detection, the study systematically evaluates various pretrained architectures and video-level embedding aggregation methods. The proposed approach significantly outperforms existing baselines across all four tasks in the ABAW-10 challenge, achieving substantial gains in robustness for affective and behavioral understanding under complex conditions while maintaining high inference efficiency.
Existing facial expression recognition datasets predominantly rely on static images, basic emotion categories, or single-label annotations, limiting their ability to capture the dynamics of facial expressions and the diversity of human perception. To address this, this work proposes Chehre—a novel, anonymized facial expression video dataset that leverages emojis as both expressive prompts and annotation primitives. The dataset employs facial motion transfer to generate synthetic videos for privacy preservation and collects multi-label emotion distributions via crowdsourcing. The study formulates a distributional facial expression recognition task and evaluates vision-language models under character-specific emoji prompts through multi-prediction assessment. Experiments reveal that state-of-the-art models achieve only 32.5% Top-1 accuracy on dominant emotion recognition, and their Spread Ratio on the distributional task remains substantially below human performance, highlighting the challenge and research potential of this new benchmark.
This study addresses the efficient assessment of consumer acceptance of branded products in supermarket or hypermarket settings by leveraging real-time analysis of facial expressions during product selection. To this end, an enhanced Harris corner detection algorithm is proposed, which significantly reduces computational time complexity while preserving high accuracy in facial expression recognition. The method optimizes the extraction of facial feature points, thereby improving the overall efficiency of the expression recognition pipeline and enhancing its suitability for real-time deployment in authentic retail environments. Experimental results demonstrate that the proposed algorithm outperforms existing approaches in corner detection speed, achieving a favorable balance between accuracy and real-time performance, thus effectively supporting public acceptance evaluation of products based on spontaneous facial expressions.