🤖 AI Summary
This paper investigates whether the layer-selection strategy for intermediate-layer matching in knowledge distillation affects student model performance. Method: Through systematic experiments, we examine the impact of matching order—forward, backward, or random—between teacher and student layers, and propose a geometric analysis framework based on inter-layer feature cosine angles to characterize the intrinsic relationship between layer alignment and distillation efficacy. Contribution/Results: We find that matching order has negligible impact on student accuracy; hand-crafted optimal layer alignment is unnecessary. Empirical validation across multiple datasets and model pairs shows that backward or random matching achieves performance within 0.3% of the best manually aligned configuration. This significantly enhances the robustness and usability of distillation methods, challenging the conventional paradigm of explicit, architecture-specific layer alignment design.
📝 Abstract
Knowledge distillation (KD) is a popular method of transferring knowledge from a large"teacher"model to a small"student"model. KD can be divided into two categories: prediction matching and intermediate-layer matching. We explore an intriguing phenomenon: layer-selection strategy does not matter (much) in intermediate-layer matching. In this paper, we show that seemingly nonsensical matching strategies such as matching the teacher's layers in reverse still result in surprisingly good student performance. We provide an interpretation for this phenomenon by examining the angles between teacher layers viewed from the student's perspective.