🤖 AI Summary
This study addresses the suboptimal optimization problem in online policy distillation, where reliance solely on positive feedback signals proves inadequate when the distributional overlap between a strong teacher and a student model is limited. To overcome this, we propose the NP-OPD framework, which introduces low-capability negative policies to generate rollouts that supply explicit negative guidance. This approach represents the first attempt to leverage negative policy rollouts for corrective supervision while preserving positive signals, requiring no modifications to the reward function. Furthermore, it integrates online policy distillation with token-level supervised learning to achieve training efficiency. Extensive experiments demonstrate that our method yields significant performance improvements across multi-scale models and reasoning tasks by effectively suppressing the generation of inferior tokens, thereby establishing a novel paradigm for knowledge distillation.
📝 Abstract
On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.