🤖 AI Summary
This work addresses the lack of mechanistic understanding, principled layer selection, and theoretically grounded feature alignment in vision transformer (ViT) knowledge distillation. We first discover that CaiT and Swin exhibit similar spectral encoding patterns, motivating a novel interpretable knowledge distillation framework based on spectral analysis—SpectralKD. Our method leverages frequency-domain spectral analysis to uncover intrinsic information transfer dynamics between teacher and student models, thereby guiding the selection of critical feature layers and enabling cross-layer spectral alignment. Evaluated on ImageNet-1k, SpectralKD improves Top-1 accuracy by 5.2% for DeiT-Tiny and 1.4% for Swin-Tiny student models. Moreover, distilled students exhibit significantly converged spectral characteristics toward their teachers, demonstrating both substantial performance gains and enhanced interpretability.
📝 Abstract
Knowledge distillation effectively reduces model complexity while improving performance, yet the underlying knowledge transfer mechanisms remain poorly understood. We propose novel spectral analysis methods and guidelines to optimize distillation, making the knowledge transfer process more interpretable. Our analysis reveals that CaiT models concentrate information in their first and last few layers, informing optimal layer selection for feature map distillation. Surprisingly, we discover that Swin Transformer and CaiT exhibit similar spectral encoding patterns despite their architectural differences, enhancing our understanding of transformer architectures and leading to improved feature map alignment strategies. Based on these insights, we introduce a simple yet effective spectral alignment method named SpectralKD. Experimental results demonstrate that following our guidelines enables SpectralKD to achieve state-of-the-art performance (DeiT-Tiny: $+5.2%$, Swin-Tiny: $+1.4%$ in ImageNet-1k Top-1 accuracy). Furthermore, through spectral analysis of student models trained with and without distillation, we show that distilled models mirror spectral patterns of their teachers, providing a new lens for interpreting knowledge distillation dynamics. Our code, pre-trained models, and experimental logs will be made publicly available.