🤖 AI Summary
This study addresses the imbalance between teacher utility and student learnability, along with privileged information leakage, in online policy self-distillation. We propose an adaptive online policy self-distillation method that dynamically modulates teacher information volume and feedback intensity based on inference directed acyclic graphs. A novel subgraph revelation mechanism evolving with student capability is introduced, combined with short-continuation probes to suppress shortcut learning. Furthermore, adaptive feedback optimization is achieved by integrating curriculum learning ranking with high-divergence tokens. Experiments demonstrate that this approach attains 72.5% Pass@8 across multiple mathematical benchmarks, outperforming the baseline by 6.7 percentage points while reducing training time by 15%.
📝 Abstract
On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.