Diagnosing On-Policy Self-Distillation for Reasoning Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability and frequent collapse of online policy self-distillation in language model reasoning, whose underlying causes remain unclear. Through controlled experiments across models ranging from 0.6B to 8B parameters, combined with token-level signal analysis and reasoning pattern alignment evaluation, this work systematically diagnoses self-distillation behaviors. It is the first to reveal the failure mechanisms of self-distillation, demonstrating that its efficacy depends on reasoning pattern alignment rather than privileged semantics, while teacher signals remain unstable and uninformative for downstream performance. Furthermore, this research establishes that the method is extremely sensitive to hyperparameters, functioning only within a narrow compatibility range. In most scenarios, it induces ineffective length inflation or behavioral collapse, confirming self-distillation as a fundamentally non-robust post-training approach.
📝 Abstract
On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privileged semantics alone. OPSD improves reasoning only in narrow compatibility regimes. Otherwise, it produces ineffective length growth, stable degradation, or behavioral collapse. Token-level analysis shows that teacher's signal is not stable and does not predict downstream performance. Based on these results, we argue that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Self-Distillation
Reasoning Language Models
Mathematical Reasoning
Behavioral Collapse
Teacher Signal
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Reasoning Language Models
Token-Level Analysis
Mathematical Reasoning
Behavioral Collapse
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.