EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition

πŸ“… 2023-09-27
πŸ›οΈ IEEE International Workshop on Multimedia Signal Processing
πŸ“ˆ Citations: 8
✨ Influential: 4
πŸ“„ PDF
πŸ€– AI Summary
To address the challenges of fine-grained emotion discrimination and model redundancy in facial expression recognition (FER), this paper proposes a lightweight and efficient framework. We first adapt ConvNeXt to FER, integrate a Spatial Transformer Network (STN) for adaptive localization of discriminative facial regions, incorporate Squeeze-and-Excitation (SE) modules to model inter-channel dependencies, and introduce a self-attention regularization mechanism to constrain feature distributions and enhance discriminative compactness. The method synergistically strengthens local-global representation learning. Evaluated on the FER2013 benchmark, our approach achieves state-of-the-art accuracy for seven basic emotion classes with significantly fewer parameters than existing methods. Ablation studies confirm the effectiveness and generalizability of both the architectural innovations and the proposed regularization strategy.

Technology Category

Computer Vision: Learning & Optimization for CVMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
πŸ“ Abstract
Facial expressions play a crucial role in human communication serving as a powerful and impactful means to express a wide range of emotions. With advancements in artificial intelligence and computer vision, deep neural networks have emerged as effective tools for facial emotion recognition. In this paper, we propose EmoNeXt, a novel deep learning framework for facial expression recognition based on an adapted ConvNeXt architecture network. We integrate a Spatial Transformer Network (STN) to focus on feature-rich regions of the face and Squeeze-and-Excitation blocks to capture channel-wise dependencies. Moreover, we introduce a self-attention regularization term, encouraging the model to generate compact feature vectors. We demonstrate the superiority of our model over existing state-of-the-art deep learning models on the FER2013 dataset regarding emotion classification accuracy.
Problem

Research questions and friction points this paper is trying to address.

Facial Expression Recognition
Accuracy Improvement
Emotion Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Transformer Networks
Squeeze-and-Excitation Blocks
Self-Attention Mechanism
πŸ”Ž Similar Papers
No similar papers found.
CESI
Yassine El Boudouri
Yassine El Boudouri
UniversitΓ© de Lille
A
Amine Bohi
CESI LINEACT Laboratory, UR 7527, Dijon, 21800, France