RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用在线策略蒸馏(OPD)作为强化学习(RL)的准备阶段,提高RL最终性能的问题。探讨了不同蒸馏方法对后续RL效果的影响。
📝 Abstract
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Policy Distillation
Training Initialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-policy distillation
reinforcement learning
performance improvement
teacher's distribution alignment
divergence objectives
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuai Dong
Fudan University; SII; JD.COM
Y
Yongfu Zhu
JD.COM
Y
Yuqi Xu
JD.COM
W
Weichu Xie
JD.COM; Peking University
L
Liuwenpu
JD.COM; Peking University
Z
Ziyue Wang
JD.COM; Peking University
K
Kaiwen Tuo
JD.COM; The Hong Kong University of Science and Technology
Congcong Wang
Congcong Wang
Norwegian University of Science and Technology
S
Siyuan Wang
The Chinese University of Hong Kong
Z
Zhongyu Wei
Fudan University; SII
J
Jiaqi Wang
JD.COM