DisFace3DNet: Explainable Facial Attractiveness Prediction via 3D Component Disentanglement

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional facial attractiveness prediction, which yields only holistic scores and fails to disentangle the contributions of independent factors such as shape, appearance, and environment. We propose a weakly supervised learning framework based on 3D component disentanglement that requires no manual component-level annotations. By jointly learning from 3D representations and image cues, the model regresses seven sub-dimensional scores from holistic ratings to quantify specific factor influences. Furthermore, weak semantic supervision and a constrained fitting algorithm are introduced to effectively decouple static from dynamic scores. Evaluated on the SCUT-FBP5500 dataset, the model achieves a Pearson correlation coefficient of 0.8904, successfully reconstructing held-out predictions and validating the roles of key components. This work establishes an interpretable, quantitative paradigm for facial attractiveness analysis.
📝 Abstract
Facial attractiveness prediction usually assigns one overall rating, leaving the roles of shape, appearance, and viewing conditions implicit. We propose DisFace3DNet, which uses 3D component disentanglement to learn seven component reference scores from overall ratings with auxiliary weak semantic supervision, without human-labeled component targets. Designated 3D representations and image cues feed jointly learned routes for identity, skin, hair, light, background, expression, and pose. A constrained fit then combines five static and two signed dynamic scores into the overall rating, exposing each component's numerical contribution and supporting component-specific comparisons across images. On SCUT-FBP5500, DisFace3DNet achieves a Pearson correlation of $0.8904\pm0.0063$ (mean $\pm$ standard deviation across five folds) with average human ratings; its component terms reconstruct every held-out prediction to numerical precision. Skin, hair, and facial shape account for the largest component-wise prediction variation. Human evaluation supports the score directions for facial shape, skin, and hair; expression agreement is weaker. DisFace3DNet thus connects overall prediction to quantitative analysis of the facial and contextual cues entering each estimate.
Problem

Research questions and friction points this paper is trying to address.

Facial attractiveness prediction
Explainability
3D component disentanglement
Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Component Disentanglement
Explainable Facial Attractiveness Prediction
Weak Semantic Supervision
Constrained Fit
Component Reference Scores
🔎 Similar Papers
No similar papers found.