π€ AI Summary
This study addresses the insufficient discriminative power of interpretable attention and the underutilization of pretrained features in fine-grained visual classification by proposing the CARE framework. Employing DINOv2 as the backbone, the method introduces a training-only auxiliary query teacher to fuse multi-layer features and optimizes the student modelβs learnable query mechanism via knowledge distillation. Additionally, diversity and sparsity regularization terms are designed to refine the attention heads. Experimental results demonstrate that CARE achieves a Top-1 accuracy of 78.5% on the CUB dataset, significantly enhancing both the faithfulness of predicted regions to class-discriminative evidence and overall recognition precision while preserving interpretability.
π Abstract
Fine-grained visual classification requires models to recognize subtle local traits while exposing the visual evidence behind their predictions. Class-specific attention pathways provide a natural basis for interpretable recognition, but their constrained prediction structure limits discriminative capacity and underuses intermediate representations from strong pretrained backbones. To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation. CARE keeps the final prediction and explanation within a class-specific attention student, while introducing a training-only auxiliary query teacher that reads selected intermediate DINOv2 layers with learnable queries. The teacher fuses multi-level representations and transfers logit-standardized class-discriminative knowledge to the student. To further refine the explanation pathway, we design diversity and sparsity terms to regularize student attention heads, reducing redundancy and encouraging compact trait localization. Experiments on CUB, Oxford-IIIT Pet, Stanford Dogs, and Stanford Cars show that CARE achieves strong classification performance under an interpretable frozen-backbone setting, reaching 78.5% Top-1 accuracy on CUB. Faithfulness analysis with insertion and deletion metrics further indicates that the top-ranked attention regions retain class-relevant evidence for explanation.