π€ AI Summary
This study addresses the scarcity of early training signals caused by sparse rewards in reinforcement learning for agents. To this end, we propose an "online warm-up" framework that first employs a teacher policy to guide the student through online policy distillation over its own trajectories, followed by a smooth transition to Reinforcement Learning with Verifiable Rewards (RLVR) and group relative optimization. Theoretical analysis demonstrates that this mechanism effectively reduces the complexity of initial reward discovery while elevating the performance upper bound. Empirical evaluations confirm that the proposed approach significantly accelerates convergence, achieving superior performance in both average and final task success rates.
π Abstract
Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.