Scaling Properties of Same-Family On-Policy Distillation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the scaling laws governing the distillation of reinforcement learning (RL)-enhanced reasoning capabilities across same-family large language models of varying sizes. Employing online policy distillation, reverse KL divergence analysis, and power-law fitting, we systematically examine scaling properties under diverse teacher-student configurations ranging from weak-to-strong setups. Our findings reveal a linear useful transfer mechanism during early training stages. Notably, compact smaller teachers outperform larger counterparts that merely match reward scores, and RL-trained compact experts effectively enhance the performance of larger student models. This work uncovers the unique advantages of smaller teachers and quantifies how teacher-student scale influences peak accuracy, providing a theoretical foundation for efficient RL knowledge distillation.
📝 Abstract
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
scaling properties
reinforcement learning
large language models
knowledge transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Scaling Laws
Weak-to-Strong Transfer
Reinforcement Learning
Reverse KL Divergence
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1