🤖 AI Summary
To address facial detail loss and emotional distortion in low-resolution facial expression recognition, this paper proposes EmoFSR—an emotion-aware facial super-resolution framework. Methodologically, it introduces an expression-preserving loss to jointly constrain expression intensity and semantic fidelity; designs an emotion-feature disentanglement module and a multi-scale perception reconstruction network; and incorporates an expression consistency constraint. Furthermore, it establishes EmoFID—the first super-resolution evaluation metric explicitly designed for emotional fidelity. Extensive experiments on CelebA, FFHQ, and Helen demonstrate that EmoFSR significantly outperforms state-of-the-art facial super-resolution (FSR) methods in both perceptual quality (PSNR/SSIM) and downstream facial expression recognition accuracy. These results empirically validate the effectiveness of co-optimizing emotional content preservation and image reconstruction quality.
📝 Abstract
Facial expression recognition (FER) systems in low-resolution settings face significant challenges in accurately identifying expressions due to the loss of fine-grained facial details. This limitation is especially problematic for applications like surveillance and mobile communications, where low image resolution is common and can compromise recognition accuracy. Traditional single-image face super-resolution (FSR) techniques, however, often fail to preserve the emotional intent of expressions, introducing distortions that obscure the original affective content. Given the inherently ill-posed nature of single-image super-resolution, a targeted approach is required to balance image quality enhancement with emotion retention. In this paper, we propose AffectSRNet, a novel emotion-aware super-resolution framework that reconstructs high-quality facial images from low-resolution inputs while maintaining the intensity and fidelity of facial expressions. Our method effectively bridges the gap between image resolution and expression accuracy by employing an expression-preserving loss function, specifically tailored for FER applications. Additionally, we introduce a new metric to assess emotion preservation in super-resolved images, providing a more nuanced evaluation of FER system performance in low-resolution scenarios. Experimental results on standard datasets, including CelebA, FFHQ, and Helen, demonstrate that AffectSRNet outperforms existing FSR approaches in both visual quality and emotion fidelity, highlighting its potential for integration into practical FER applications. This work not only improves image clarity but also ensures that emotion-driven applications retain their core functionality in suboptimal resolution environments, paving the way for broader adoption in FER systems.