š¤ AI Summary
This study addresses the problem of vanishing probability mass in Direct Online Policy Distillation (Direct-OPD), which renders supervision signals ineffective. To overcome this, we propose S²D-OPD, a method grounded in theoretical analysis that introduces a selective supervision mechanism. Specifically, it employs Jensen-Shannon divergence to identify high-value states and applies token-level masking to suppress ineffective supervision associated with low-divergence tokens, retaining only the top 10% of critical states for efficient training. Experimental results demonstrate that S²D-OPD outperforms dense supervision baselines in accuracy across most scenarios on the AIME and HMMT benchmarks. Notably, this improvement is achieved without introducing additional computational overhead, thereby realizing highly efficient policy distillation.
š Abstract
Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S$^2$D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S$^2$D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes. Our code is available at https://anonymous.4open.science/r/S2D-OPD-8868.