🤖 AI Summary
This study addresses the accuracy and bias challenges in zero-shot gender estimation from full-face and periocular images. Methodologically, it evaluates the training-free gender recognition capability of the CLIP vision-language model, revealing a significant male bias mechanism in zero-shot predictions for the periocular region. To mitigate this, a threshold alignment strategy is proposed to optimize the decision boundary and eliminate such bias. Experimental results demonstrate that full-face classification achieves an accuracy of 95.54%, while periocular performance reaches 85.29% after applying threshold alignment. This work validates the inherent difficulty of periocular gender estimation and confirms that the proposed alignment strategy effectively enhances fairness and robustness in zero-shot scenarios.
📝 Abstract
We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation