DisFace3DNet: Explainable Facial Attractiveness Prediction via 3D Component Disentanglement
This study addresses the limitation of conventional facial attractiveness prediction, which yields only holistic scores and fails to disentangle the contributions of independent factors such as shape, appearance, and environment. We propose a weakly supervised learning framework based on 3D component disentanglement that requires no manual component-level annotations. By jointly learning from 3D representations and image cues, the model regresses seven sub-dimensional scores from holistic ratings to quantify specific factor influences. Furthermore, weak semantic supervision and a constrained fitting algorithm are introduced to effectively decouple static from dynamic scores. Evaluated on the SCUT-FBP5500 dataset, the model achieves a Pearson correlation coefficient of 0.8904, successfully reconstructing held-out predictions and validating the roles of key components. This work establishes an interpretable, quantitative paradigm for facial attractiveness analysis.