An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sample efficiency and lack of diversity in policy distillation for large language model (LLM) reasoning by proposing the LSPD framework. Methodologically, it reformulates online policy distillation from a reinforcement learning perspective, establishing a theoretical connection between the reverse KL objective and KL-regularized policy optimization. By integrating least-squares policy distillation, optimistic value exploration, and offline data reuse mechanisms, the framework enables efficient offline learning while preserving policy diversity. Experimental results demonstrate that LSPD achieves an average improvement of 1.59 points on mathematical reasoning benchmarks, attaining comparable performance using only 25% of the training batches, and significantly enhancing diversity as measured by the Pass@k metric.
📝 Abstract
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Sample efficiency
Policy diversity
LLM reasoning
Rollout efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Distillation
Reinforcement Learning
Least-Square Policy Distillation
Off-Policy Data Reuse
Optimistic Exploration
🔎 Similar Papers
No similar papers found.