🤖 AI Summary
This study addresses the conflict between low probability and high confidence in teacher distributions during knowledge distillation for small models, which often degrades generalization and induces catastrophic forgetting. To mitigate this, we propose an entropy-aware mixing mechanism that introduces a novel token-level predictive entropy-gated strategy for convex and geometric distribution interpolation, dynamically fusing student-teacher output distributions. We further tailor entropy scheduling schemes for both offline and online distillation settings. Combined with speculative decoding, supervised fine-tuning (SFT), and forward KL divergence, our approach enables efficient training. The proposed method effectively stabilizes the distillation process, substantially enhances students' mathematical reasoning capabilities, and better preserves their general language understanding and generation performance.
📝 Abstract
Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e., in a low-entropy state). To mitigate the degradation in generalisation and catastrophic forgetting caused by these conflicts, we propose \textbf{Entropy-Aware Mixing}: a dynamic per-token interpolation of the student and teacher distributions, gated by the student's predictive entropy. We implement both convex and geometric interpolations for both offline trace generation (via speculative decoding, then SFT) and on-policy forward-KL distillation. Our results show that entropy-aware mixing stabilises distillation, improving in-distribution and out-of-distribution math reasoning while better preserving general capabilities than fixed-teacher supervision. Nonetheless, the optimal entropy schedule depends on the training source, with offline-generated traces favouring concave schedules (greater overall teacher influence) and on-policy training favouring linear or convex schedules (teacher concentrated in high-entropy states).