🤖 AI Summary
This work addresses the long-standing lack of theoretical guidance in selecting the temperature parameter for knowledge distillation, a choice often made empirically or via grid search, frequently yielding suboptimal student performance. For the first time, the study systematically investigates the interaction between temperature and key training factors—such as optimizer choice and whether the teacher model is pretrained or fine-tuned—and identifies representative scenarios that critically influence optimal temperature selection. Building upon a classification distillation framework and a comprehensive cross-factor experimental design, the authors propose a context-aware temperature selection strategy. This approach moves beyond conventional fixed or heuristic tuning paradigms, offering practical guidelines tailored to diverse training configurations and consistently achieving significant improvements in student model performance.
📝 Abstract
A central idea of knowledge distillation is to expose relational structure embedded in the teacher's weights for the student to learn, which is often facilitated using a temperature parameter. Despite its widespread use, there remains limited understanding on how to select an appropriate temperature value, or how this value depends on other training elements such as optimizer, teacher pretraining/finetuning, etc. In practice, temperature is commonly chosen via grid search or by adopting values from prior work, which can be time-consuming or may lead to suboptimal student performance when training setups differ. In this work, we posit that temperature is closely linked to these training components and present a unified study that systematically examines such interactions. From analyzing these cross-connections, we identify and present common situations that have a pronounced impact on temperature selection, providing valuable guidance for practitioners employing knowledge distillation in their work.