π€ AI Summary
This study addresses the synergistic failures arising from the isolated design of Supervised Fine-Tuning (SFT), Reinforcement Learning with Verifiable Rewards (RLVR), and Online Policy Distillation (OPD) stages during LLM post-training. Through systematic controlled experiments on mathematical reasoning tasks using Qwen3 models, we reveal the interdependent mechanisms among these stages regarding distillation efficacy and initialization, and propose an OPD optimization strategy grounded in teacher-student compatibility. Our findings demonstrate that SFT warm-up combined with teacher RLVR adaptation significantly enhances downstream performance. Compared to standalone SFT, this combination provides superior initialization for subsequent RLVR, with advantages amplifying as compute scales, ultimately improving OPD accuracy by 50%.
π Abstract
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves the current model can make the next stage less effective. Through controlled experiments with Qwen3 models on math and science reasoning, we first characterize OPD across nine student-teacher pairs spanning 2x to 53x parameter ratios and show that OPD effectiveness depends on student-teacher compatibility rather than teacher scale alone. The surrounding stages of OPD reshape this compatibility in three ways: (1) A brief SFT warm-up improves subsequent OPD, while an RLVR-strengthened student regresses under distillation from the same teacher. (2) Adapting the teacher with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Following these two interventions, we find that combining teacher adaptation and student warm-up alone raise average OPD accuracy from 29.2\% to 43.8\% (50\% relative improvement) after the same number of distillation steps, with additional preparatory training. (3) At comparable accuracy, OPD leaves a stronger initialization for downstream RLVR than SFT, with a gap that widens as RL compute scales. Our results suggest that each post-training stage should be chosen not only for the capability it adds, but for the learning interface it creates for the next stage.