🤖 AI Summary
This study addresses the instability and frequent collapse of online policy self-distillation in language model reasoning, whose underlying causes remain unclear. Through controlled experiments across models ranging from 0.6B to 8B parameters, combined with token-level signal analysis and reasoning pattern alignment evaluation, this work systematically diagnoses self-distillation behaviors. It is the first to reveal the failure mechanisms of self-distillation, demonstrating that its efficacy depends on reasoning pattern alignment rather than privileged semantics, while teacher signals remain unstable and uninformative for downstream performance. Furthermore, this research establishes that the method is extremely sensitive to hyperparameters, functioning only within a narrow compatibility range. In most scenarios, it induces ineffective length inflation or behavioral collapse, confirming self-distillation as a fundamentally non-robust post-training approach.
📝 Abstract
On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privileged semantics alone. OPSD improves reasoning only in narrow compatibility regimes. Otherwise, it produces ineffective length growth, stable degradation, or behavioral collapse. Token-level analysis shows that teacher's signal is not stable and does not predict downstream performance. Based on these results, we argue that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method.