Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imbalance between teacher utility and student learnability, along with privileged information leakage, in online policy self-distillation. We propose an adaptive online policy self-distillation method that dynamically modulates teacher information volume and feedback intensity based on inference directed acyclic graphs. A novel subgraph revelation mechanism evolving with student capability is introduced, combined with short-continuation probes to suppress shortcut learning. Furthermore, adaptive feedback optimization is achieved by integrating curriculum learning ranking with high-divergence tokens. Experiments demonstrate that this approach attains 72.5% Pass@8 across multiple mathematical benchmarks, outperforming the baseline by 6.7 percentage points while reducing training time by 15%.
📝 Abstract
On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.
Problem

Research questions and friction points this paper is trying to address.

On-policy self-distillation
Mathematical reasoning
Privileged information leakage
Teacher-student gap
Shortcut learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Reasoning DAG
Adaptive Privileged Information
Mathematical Reasoning
Shortcut Mitigation