Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

šŸ“… 2026-09-24
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study addresses the problem of vanishing probability mass in Direct Online Policy Distillation (Direct-OPD), which renders supervision signals ineffective. To overcome this, we propose S²D-OPD, a method grounded in theoretical analysis that introduces a selective supervision mechanism. Specifically, it employs Jensen-Shannon divergence to identify high-value states and applies token-level masking to suppress ineffective supervision associated with low-divergence tokens, retaining only the top 10% of critical states for efficient training. Experimental results demonstrate that S²D-OPD outperforms dense supervision baselines in accuracy across most scenarios on the AIME and HMMT benchmarks. Notably, this improvement is achieved without introducing additional computational overhead, thereby realizing highly efficient policy distillation.
šŸ“ Abstract
Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S$^2$D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S$^2$D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes. Our code is available at https://anonymous.4open.science/r/S2D-OPD-8868.
Problem

Research questions and friction points this paper is trying to address.

Direct On-Policy Distillation
Selective Supervision
Token-level Log-ratio
Jensen-Shannon Divergence
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct On-Policy Distillation
Selective Supervision
Jensen-Shannon Divergence
Knowledge Distillation
Reinforcement Learning
šŸ”Ž Similar Papers
No similar papers found.
šŸ’¼ Related Jobs
No related jobs found.
Yibo Zhao
Yibo Zhao
Student of Data Science, East China Normal University
AI
Z
Zixuan Yang
School of Data Science and Engineering, East China Normal University; Hugging Face
Y
Yunshi Lan
School of Data Science and Engineering, East China Normal University
Xiang Li
Xiang Li
East China Normal University
Data miningmachine learning