From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the scarcity of early training signals caused by sparse rewards in reinforcement learning for agents. To this end, we propose an "online warm-up" framework that first employs a teacher policy to guide the student through online policy distillation over its own trajectories, followed by a smooth transition to Reinforcement Learning with Verifiable Rewards (RLVR) and group relative optimization. Theoretical analysis demonstrates that this mechanism effectively reduces the complexity of initial reward discovery while elevating the performance upper bound. Empirical evaluations confirm that the proposed approach significantly accelerates convergence, achieving superior performance in both average and final task success rates.
πŸ“ Abstract
Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Reward Discovery
Language Model Agents
Sparse Rewards
RLVR
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Warmup
Reinforcement Learning with Verifiable Reward (RLVR)
Reverse-KL Distillation
Reward Discovery
Agentic RL