E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of online self-distillation, where non-transferable supervision signals and entropy overshoot induced by forward KL divergence constrain model reasoning capabilities. To overcome these challenges, this work proposes an exemplar-guided teaching strategy coupled with an entropy-aware dynamic correction mechanism that optimizes the transfer of reasoning patterns while calibrating output distributions. Notably, the approach achieves lightweight online self-distillation without requiring additional networks or extra forward passes. Empirical evaluations demonstrate that the proposed method effectively suppresses entropy overshoot and enhances generalization, yielding an average accuracy improvement of 4.3 points on mathematical reasoning benchmarks. Furthermore, it exhibits significantly superior out-of-domain performance compared to baseline models.
📝 Abstract
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E$^2$-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E$^2$-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E$^2$-OPSD remains simple, requiring no additional forward passes or networks.
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy self-distillation
Entropy overshoot
Exemplar-guided teaching
Entropy-aware distillation
Math reasoning
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1