🤖 AI Summary
This study addresses the inherent trade-off in agent reinforcement learning between training engine idleness under synchronous execution and policy staleness under asynchronous execution. To reconcile this dilemma, we propose a fine-grained gradient flow mechanism that guarantees zero policy staleness. By processing single-trajectory or single-round data on the fly within Group Relative Policy Optimization (GRPO) and Online Policy Distillation (OPD), our method is theoretically proven to yield updates mathematically identical to those of batch synchronous training. Experimental results demonstrate that the proposed mechanism accelerates training by up to 1.9× compared to synchronous baselines while improving performance by up to 2.47 percentage points over asynchronous methods under a fixed computational budget, thereby successfully unifying training efficiency with optimization accuracy.
📝 Abstract
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to $1.9 \times$ faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to $2.47$ percentage points.