π€ AI Summary
This study addresses the challenge of effectively distilling reasoning capabilities from reinforcement learning-enhanced teacher models into smaller student models. To this end, we propose TS-OPD, a method that hierarchically filters problems based on joint teacher-student success rates, dynamically routing them to either forward or reverse KL divergence objectives while introducing an entropy braking mechanism to preserve sampling coverage. The core innovation lies in revealing the structured nature of knowledge transfer, adaptively aligning supervision direction and token budgets with the studentβs observational capacity, thereby overcoming paradigms limited solely by teacher strength. Experiments demonstrate that our approach significantly improves macro-average accuracy on mathematical reasoning tasks, validating the critical role of the proposed routing and gating mechanisms.
π Abstract
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.