🤖 AI Summary
Existing face recognition models exhibit insufficient robustness under Out-of-Distribution (OOD) conditions, particularly degrading significantly when confronted with real-world image degradations and appearance variations. To address this, we introduce OODFace—the first OOD-robustness benchmark specifically designed for face verification—comprising 30 common degradation types (e.g., noise, blur, compression) and appearance variations (e.g., pose, illumination, masks, makeup), and establishing three evaluation splits: LFW-C/V, CFP-FP-C/V, and YTF-C/V. We propose the first dual-dimensional (degradation + appearance) modeling framework for facial OOD challenges, along with a scalable, unified evaluation toolkit. Extensive experiments reveal that 19 open-source models and 3 commercial APIs suffer over 40% accuracy drops under occlusion, illumination shifts, and mask perturbations. We further validate the efficacy of vision-language model–assisted analysis and physical-world experiments. All code, data, and evaluation tools are publicly released.
📝 Abstract
With the rise of deep learning, facial recognition technology has seen extensive research and rapid development. Although facial recognition is considered a mature technology, we find that existing open-source models and commercial algorithms lack robustness in certain complex Out-of-Distribution (OOD) scenarios, raising concerns about the reliability of these systems. In this paper, we introduce OODFace, which explores the OOD challenges faced by facial recognition models from two perspectives: common corruptions and appearance variations. We systematically design 30 OOD scenarios across 9 major categories tailored for facial recognition. By simulating these challenges on public datasets, we establish three robustness benchmarks: LFW-C/V, CFP-FP-C/V, and YTF-C/V. We then conduct extensive experiments on 19 facial recognition models and 3 commercial APIs, along with extended physical experiments on face masks to assess their robustness. Next, we explore potential solutions from two perspectives: defense strategies and Vision-Language Models (VLMs). Based on the results, we draw several key insights, highlighting the vulnerability of facial recognition systems to OOD data and suggesting possible solutions. Additionally, we offer a unified toolkit that includes all corruption and variation types, easily extendable to other datasets. We hope that our benchmarks and findings can provide guidance for future improvements in facial recognition model robustness.