🤖 AI Summary
This work addresses the limited ability of convolutional neural networks (CNNs) to effectively model directional features, which constrains periocular recognition performance. To overcome this, the authors propose replacing conventional grayscale inputs with complex structural tensors that encode compact directional information along with its confidence. For the first time, an explicit orientation prior—inspired by mammalian visual mechanisms—is integrated into the front end of CNNs. The approach combines a lightweight complex-valued convolutional network with six mainstream CNN architectures and is evaluated under both closed-set and open-set biometric protocols. Experiments on the Cross-Eyed and PolyU datasets demonstrate substantial improvements, achieving 5%–26% lower equal error rates (EER) compared to full-scale state-of-the-art models, while simultaneously enhancing model interpretability and suitability for edge deployment.
📝 Abstract
Our study provides evidence that CNNs struggle to extract orientation features effectively. We show that using the Complex Structure Tensor, which contains compact orientation features with certainties, as input to CNNs consistently improves identification accuracy compared to grayscale inputs alone. Experiments also demonstrated that our inputs, provided by mini-complex convnets, combined with reduced CNN sizes, outperformed full-fledged, prevailing CNN architectures. This suggests that the upfront use of orientation features in CNNs, a strategy seen in mammalian vision, not only mitigates their limitations but also enhances their explainability and relevance to thin-clients. Experiments were conducted on publicly available datasets comprising periocular images (Cross-Eyed and PolyU) for biometric identification and verification in both Close-World and Open-World Scenarios using six CNN architectures. Our experiments on the Cross-Eyed and PolyU datasets yield a 5-26% reduction in EER, providing strong empirical evidence that explicit orientation priors mitigate CNN representational limits in Open-World and Close-World scenarios.