Score
Designs, trains, and evaluates models and end-to-end pipelines that perform identity verification or categorical identity classification from biometric inputs (commonly face images), including feature extraction, loss functions, and training procedures. Builds and analyzes system components for thresholding, evaluation metrics, robustness to variation, and fairness/privacy considerations for deployed face-recognition systems.
This study addresses the lack of compliant and practical threshold calibration methods for face verification systems in border control under extremely low false match rates (FMRs), hindered by legal and privacy constraints on real-world data acquisition. It presents the first systematic evaluation of synthetic facial data for calibrating identity-to-live face verification thresholds in high-security scenarios. Through score distribution alignment, cross-domain threshold transfer, and adversarial morphing attack testing, the work demonstrates that synthetic data can approximate real-data calibration behavior under controlled conditions. However, under uncontrolled settings, performance degrades significantly due to tail distribution mismatches, introducing notable security vulnerabilities. The findings reveal that calibration efficacy is highly dataset-dependent, highlighting the limited generalizability of synthetic data in real-world deployment contexts.
This study addresses the challenge that complex backgrounds in unconstrained scenarios—such as airport border control—degrade face recognition accuracy and impair the detection of presentation attacks. The authors systematically evaluate the impact of multiple face segmentation methods on four representative recognition models and three attack detection techniques through comprehensive experiments on datasets encompassing both controlled and unconstrained imagery. For the first time, they comprehensively demonstrate the dual role of background removal in simultaneously influencing recognition performance and security mechanisms, showing its significant effects on image quality, identification accuracy, and attack detectability. These findings provide empirical grounding and practical guidance for preprocessing strategies in real-world biometric systems, effectively bridging the critical gap between deployment feasibility and system reliability.
This work proposes a low-cost, dual-modal biometric authentication system that leverages only a standard camera and microphone, addressing the high cost and limited practicality of traditional identity verification systems reliant on specialized hardware. The approach employs a two-stage cascaded mechanism: an initial screening stage utilizes a pruned VGG-16 network integrated with MTCNN for face recognition, followed by a secondary verification stage that applies a CNN-based speaker verification model for matched identities. This design effectively balances computational efficiency and robustness. Experimental results demonstrate that the system achieves a face recognition accuracy of 95.1% and a voice verification accuracy of 98.9%, with an equal error rate (EER) as low as 3.456%, significantly enhancing both security and usability in scenarios devoid of dedicated authentication hardware.
To address security vulnerabilities in biometric systems caused by face morphing attacks, this paper proposes an end-to-end robustness enhancement method that enables models to actively reject forged faces while preserving high identity verification accuracy. The core method integrates deep metric learning, geometric constraints in the embedding space, and differentiable training. We introduce TetraLoss—the first quartet-based contrastive loss—explicitly enlarging the embedding distance between morphed faces and their source identities, thereby jointly optimizing recognition performance and spoof detection capability. Evaluated on multiple morphing benchmarks, our approach reduces attack success rates by over 60%, significantly outperforming state-of-the-art reference-free detectors and joint-learning baselines. Crucially, it achieves this without any degradation in original face verification accuracy—maintaining zero loss in genuine-user recognition performance.
Face recognition systems risk leaking training data privacy through membership inference attacks. Method: This paper proposes MINT, the first large-scale privacy auditing framework for facial images, featuring a dual-architecture discriminative model (MLP + CNN) that captures activation pattern disparities between member and non-member samples. MINT establishes a multi-source experimental framework—spanning diverse databases and state-of-the-art (SOTA) face recognition models—to enable cross-dataset and cross-model evaluation. Contribution/Results: Evaluated on a real-world dataset of 22 million face images, MINT achieves up to 90% membership inference accuracy, substantially outperforming existing baselines. It is the first study to empirically demonstrate the feasibility of membership inference in large-scale, realistic facial recognition settings. By providing a deployable, scalable privacy assessment tool, MINT advances compliance auditing for training data used in large language models and vision models.
研究针对图像生成器的隐私保护问题,通过比较四种审计方法(GaussMech、KDE-LR、MMD-TV、ROC-HT)来评估身份级别的差分隐私参数ε。
Facial recognition suffers from poor interpretability, demographic bias, privacy risks, and limited robustness to aging, pose, illumination, occlusion, and expression—exacerbated by privacy regulations that degrade real-data quality. This work systematically evaluates three synthetic face generation paradigms—diffusion models, GANs, and 3D modeling—across eight benchmark datasets for facial recognition tasks. It presents the first comprehensive empirical analysis of multi-paradigm synthetic data in bias mitigation, real-data substitution capability, and model generalization enhancement. Results demonstrate that synthetic data effectively captures facial variability, significantly improving model robustness and mitigating gender and racial biases; in certain scenarios, performance approaches that of real data. However, overall accuracy and true positive rates remain marginally lower, revealing current generative methods’ practical limits. This study fills a critical gap by providing the first cross-method, cross-dataset systematic evaluation of synthetic data in facial recognition.
This work addresses the challenges in biometric systems arising from stringent privacy regulations and limited access to real facial data, as well as the difficulty of existing generative models in simultaneously achieving high realism, diversity, and identity preservation. To this end, we propose a novel synthetic face generation pipeline that integrates StyleCLIP for semantic control, HyperStyle for high-fidelity reconstruction, InterfaceGAN for attribute editing, and diffusion models for generative capacity into an end-to-end controllable framework. This approach enables high-quality intra-class variation and inter-class distinction while preserving identity consistency. Experimental results demonstrate that the synthesized dataset achieves image quality and diversity comparable to real data under ArcFace verification, effectively supporting biometric system evaluation and showing strong potential to partially replace real-world data.
This study addresses the dual challenge in face verification systems: defending against presentation attacks (e.g., photos, videos) while maintaining robustness to legitimate appearance variations such as accessories, lighting, and pose changes. The authors propose a unified framework that, for the first time, integrates classical handcrafted features—including PCA, LBP, HOG, SURF, and Harris corner detectors—within five fusion strategies: Product Matching (PM), Linear Product Matching (LPM), Hierarchical Product Matching (HPM), Sum Matching (SM), and Hybrid Matching (HM). These strategies jointly optimize spoof detection and recognition robustness across preprocessing and classification stages. Experimental results demonstrate that HPM achieves 94.59% accuracy under mixed spoofing attacks, 81.5%–93.2% under lighting and pose variations, and 91.67% anti-spoofing performance; LPM yields the best anti-spoofing rate (93.2%) but exhibits weaker pose robustness. The work further reveals a quantifiable sensitivity–robustness trade-off between spoof detection and appearance invariance.
This study investigates the geometric separability between embeddings of training (member) and non-training (non-member) identities in hyperspherical embedding spaces of face recognition models, along with its underlying factors. Through a factorial experimental design, the authors train 180 models varying across four dimensions: IResNet backbone architecture, loss function, training duration, and number of training identities. Four clustering-based geometric statistics are computed to quantify distributional differences between member and non-member embeddings. The work reveals, for the first time, that the number of training identities is the dominant factor governing member/non-member separability, and that out-of-domain non-member data can exaggerate membership signals. Furthermore, fusing multiple geometric statistics significantly enhances membership inference performance. Experiments across nine benchmarks consistently show a monotonic negative correlation between the number of training identities and membership signal strength, while the effects of backbone architecture and loss function are comparatively weak.