🤖 AI Summary
This work addresses the pervasive age bias in deep learning–based facial expression recognition (FER), particularly the unfair performance degradation observed for older adults. We propose three lightweight mitigation strategies: (1) multi-task learning jointly predicting expression and age, (2) multimodal input integrating RGB and geometric facial features, and (3) an age-weighted loss function. An age-aware FER model is trained on AffectNet augmented with automatically estimated age labels. To uncover age-related decision biases, we employ eXplainable AI (XAI) techniques—specifically attention visualization—to analyze model behavior across age groups. Experiments on balanced benchmarks demonstrate significant accuracy gains for elderly subjects, especially for ambiguous expressions such as “neutral,” “sad,” and “angry.” Attention heatmaps further confirm enhanced focus on physiologically plausible facial regions. Crucially, this study provides the first systematic empirical validation of the critical role played by approximate demographic annotations—here, age—in ensuring fairness in affective computing.
📝 Abstract
Facial Expression Recognition (FER) systems based on deep learning have achieved impressive performance in recent years. However, these models often exhibit demographic biases, particularly with respect to age, which can compromise their fairness and reliability. In this work, we present a comprehensive study of age-related bias in deep FER models, with a particular focus on the elderly population. We first investigate whether recognition performance varies across age groups, which expressions are most affected, and whether model attention differs depending on age. Using Explainable AI (XAI) techniques, we identify systematic disparities in expression recognition and attention patterns, especially for "neutral", "sadness", and "anger" in elderly individuals. Based on these findings, we propose and evaluate three bias mitigation strategies: Multi-task Learning, Multi-modal Input, and Age-weighted Loss. Our models are trained on a large-scale dataset, AffectNet, with automatically estimated age labels and validated on balanced benchmark datasets that include underrepresented age groups. Results show consistent improvements in recognition accuracy for elderly individuals, particularly for the most error-prone expressions. Saliency heatmap analysis reveals that models trained with age-aware strategies attend to more relevant facial regions for each age group, helping to explain the observed improvements. These findings suggest that age-related bias in FER can be effectively mitigated using simple training modifications, and that even approximate demographic labels can be valuable for promoting fairness in large-scale affective computing systems.