Outcome-Guided On-Policy Self-Distillation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the noise introduction and training instability caused by dense supervision in online self-distillation by proposing the OG-OPSD method. Overcoming the limitations of fixed divergence, this approach adaptively regulates divergence targets and distillation positions by integrating outcome correctness signals with cumulative average teacher entropy, thereby eliminating the need for additional hyperparameter balancing. Technically, it synthesizes online self-distillation, reinforcement learning advantage function analysis, and dynamic weight adjustment mechanisms. Experiments on the Qwen3 model series demonstrate that the proposed method significantly enhances performance in mathematical reasoning, multimodal understanding, and out-of-distribution tasks, comprehensively outperforming existing strong baselines.
📝 Abstract
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
Problem

Research questions and friction points this paper is trying to address.

On-policy self-distillation
training instability
outcome correctness
teacher entropy
noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Outcome-Guided
Teacher Entropy
Dynamic Divergence Objective
Reinforcement Learning with Verifiable Rewards
🔎 Similar Papers
No similar papers found.
Z
ZheXu Wang
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China
M
Mao-Lin Luo
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China
Y
Yankun Hong
Huawei Noah’s Ark Lab
Z
Zi-Hao Zhou
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China
B
Bo Ye
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China; Zhongguancun Academy
Jian Zhao
Jian Zhao
Zhongguancun Institute of Artificial Intelligence
Reinforcement LearningMulti-Agent System
X
Xialiang Tong
Huawei Noah’s Ark Lab
Min-Ling Zhang
Min-Ling Zhang
Professor, School of Computer Science and Engineering, Southeast University, China
Artificial IntelligenceMachine LearningData Mining
Tong Wei
Tong Wei
Southeast University
Machine Learning