🤖 AI Summary
This study addresses the noise introduction and training instability caused by dense supervision in online self-distillation by proposing the OG-OPSD method. Overcoming the limitations of fixed divergence, this approach adaptively regulates divergence targets and distillation positions by integrating outcome correctness signals with cumulative average teacher entropy, thereby eliminating the need for additional hyperparameter balancing. Technically, it synthesizes online self-distillation, reinforcement learning advantage function analysis, and dynamic weight adjustment mechanisms. Experiments on the Qwen3 model series demonstrate that the proposed method significantly enhances performance in mathematical reasoning, multimodal understanding, and out-of-distribution tasks, comprehensively outperforming existing strong baselines.
📝 Abstract
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.