A Unified Revisit of Temperature in Classification-Based Knowledge Distillation

📅 2026-03-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the long-standing lack of theoretical guidance in selecting the temperature parameter for knowledge distillation, a choice often made empirically or via grid search, frequently yielding suboptimal student performance. For the first time, the study systematically investigates the interaction between temperature and key training factors—such as optimizer choice and whether the teacher model is pretrained or fine-tuned—and identifies representative scenarios that critically influence optimal temperature selection. Building upon a classification distillation framework and a comprehensive cross-factor experimental design, the authors propose a context-aware temperature selection strategy. This approach moves beyond conventional fixed or heuristic tuning paradigms, offering practical guidelines tailored to diverse training configurations and consistently achieving significant improvements in student model performance.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
A central idea of knowledge distillation is to expose relational structure embedded in the teacher's weights for the student to learn, which is often facilitated using a temperature parameter. Despite its widespread use, there remains limited understanding on how to select an appropriate temperature value, or how this value depends on other training elements such as optimizer, teacher pretraining/finetuning, etc. In practice, temperature is commonly chosen via grid search or by adopting values from prior work, which can be time-consuming or may lead to suboptimal student performance when training setups differ. In this work, we posit that temperature is closely linked to these training components and present a unified study that systematically examines such interactions. From analyzing these cross-connections, we identify and present common situations that have a pronounced impact on temperature selection, providing valuable guidance for practitioners employing knowledge distillation in their work.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
temperature selection
training components
student performance
hyperparameter tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge distillation
temperature scaling
teacher-student learning
hyperparameter analysis
training dynamics
🔎 Similar Papers