SpectralKD: Understanding and Optimizing Vision Transformer Distillation through Spectral Analysis

📅 2024-12-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of mechanistic understanding, principled layer selection, and theoretically grounded feature alignment in vision transformer (ViT) knowledge distillation. We first discover that CaiT and Swin exhibit similar spectral encoding patterns, motivating a novel interpretable knowledge distillation framework based on spectral analysis—SpectralKD. Our method leverages frequency-domain spectral analysis to uncover intrinsic information transfer dynamics between teacher and student models, thereby guiding the selection of critical feature layers and enabling cross-layer spectral alignment. Evaluated on ImageNet-1k, SpectralKD improves Top-1 accuracy by 5.2% for DeiT-Tiny and 1.4% for Swin-Tiny student models. Moreover, distilled students exhibit significantly converged spectral characteristics toward their teachers, demonstrating both substantial performance gains and enhanced interpretability.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Feature Construction/ReformulationKnowledge Representation and Reasoning: Knowledge Engineering

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web query analysis, representation and understandingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Knowledge distillation effectively reduces model complexity while improving performance, yet the underlying knowledge transfer mechanisms remain poorly understood. We propose novel spectral analysis methods and guidelines to optimize distillation, making the knowledge transfer process more interpretable. Our analysis reveals that CaiT models concentrate information in their first and last few layers, informing optimal layer selection for feature map distillation. Surprisingly, we discover that Swin Transformer and CaiT exhibit similar spectral encoding patterns despite their architectural differences, enhancing our understanding of transformer architectures and leading to improved feature map alignment strategies. Based on these insights, we introduce a simple yet effective spectral alignment method named SpectralKD. Experimental results demonstrate that following our guidelines enables SpectralKD to achieve state-of-the-art performance (DeiT-Tiny: $+5.2%$, Swin-Tiny: $+1.4%$ in ImageNet-1k Top-1 accuracy). Furthermore, through spectral analysis of student models trained with and without distillation, we show that distilled models mirror spectral patterns of their teachers, providing a new lens for interpreting knowledge distillation dynamics. Our code, pre-trained models, and experimental logs will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
Image Processing
Complex Model Simplification
Innovation

Methods, ideas, or system contributions that make the work stand out.

SpectralKD
Knowledge Distillation
Transformer Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhejiang University
H
Huiyuan Tian
College of Computer Science and Technology, Zhejiang University, NO. 38 Zheda Road, Xihu District, Hangzhou 310027, China
B
Bonan Xu
School of Aeronautics and Astronautics, Zhejiang University, NO. 38 Zheda Road, Xihu District, Hangzhou 310027, China
Shijian Li
Shijian Li
zhejiang university
pervasive computinghuman computer interactionartificial intelligence
Gang Pan
Gang Pan
Tianjin University
Computer visionMultimodalAI