๐ค AI Summary
This study addresses the limitation of traditional knowledge distillation, where blindly imitating the teacherโs probability distribution compromises the student modelโs inherent discriminative capacity. To overcome this, we propose Task-Preserving Knowledge Distillation (TPKD). Grounded in a theoretical analysis of KL divergence, TPKD precisely decouples full imitation from conditional learning, facilitating superior knowledge transfer by retaining label gradients while rectifying conditional gradients. Theoretically, we demonstrate that the rectified conditional direction preserves over half of the first-order descent benefits. Experimental results show that TPKD achieves accuracies of 88.05% and 93.81% on CIFAR-100 and CLINC150, respectively, significantly outperforming standard distillation methods.
๐ Abstract
Knowledge distillation transfers knowledge by encouraging a student to match a teacher's predicted class probabilities. These probabilities express not only confidence in the correct class, but also relations among incorrect alternatives. Yet closer imitation does not necessarily yield a better student. A student may already distinguish the correct class more sharply than its teacher, so further imitation can require giving back discrimination it has acquired. Our main result is an exact separation between full imitation and conditional learning. When the correct class's score advantage over each alternative must be preserved, full teacher-to-student KL minimization is blocked exactly when the student assigns no more probability than the teacher to every incorrect class. Crucially, the teacher's relative probabilities among incorrect classes remain fully learnable. We characterize the exact price of this transfer: a minimum increase in correct-class log-odds that compensates for the largest conditional-probability mismatch. Label fitting and conditional matching can therefore be completed even as full teacher KL diverges. This separation motivates task-preserving knowledge distillation (TPKD), which keeps the label gradient intact and minimally corrects the conditional gradient so that its output update preserves the label step's gains against every incorrect alternative. The corrected conditional direction retains more than half of the original first-order conditional descent at the same step size, with a tight bound. For a fixed positive conditional target and sufficiently small constant output steps, label and conditional errors vanish together. Experiments trace this learning from exact head updates to ordinary network training. TPKD reaches 88.05% accuracy on CIFAR-100 and 93.81% on CLINC150, improving over standard distillation by 0.47 and 0.35 percentage points across three seeds.