🤖 AI Summary
This study addresses the training throughput bottlenecks caused by uneven trajectory completion in agent reinforcement learning, as well as resource waste stemming from static memory allocation in tool sandboxes. To overcome these challenges, we propose a fully decoupled architecture. Methodologically, we design a length-prediction-based, priority-aware action scheduler that accelerates the completion of critical sample groups to unlock training steps. Furthermore, we construct a template-keyed page-sharing pool that leverages copy-on-write and write-protected aliasing mechanisms to achieve efficient isolation and reuse of environment memory. Experimental results demonstrate that the proposed system achieves up to 4.24× end-to-end training speedup over state-of-the-art baselines while reducing environment resource overhead by 89%.
📝 Abstract
Agentic Reinforcement Learning (RL) trains LLM agents through multi-turn interactions with external tool environments. Its multi-turn nature exposes two system-level bottlenecks unaddressed by existing agentic RL frameworks. First, end-to-end training throughput is constrained by the slowest trajectories to complete, yet optimizing per-GPU utilization alone scatters rollout progress across many groups, delaying the completion of enough groups to unblock the next training step. Second, tool sandboxes are statically over-provisioned by their declared memory ceilings, leaving most physical memory stranded while replicating near-identical state across sandboxes launched from the same prompt. We present VenusRL, a fully disaggregated agentic RL system that addresses both bottlenecks. VenusRL's priority-aware action scheduler uses length-prediction heuristics to identify sample groups whose completion is most likely to unblock the next training step, and pushes them ahead of others across batch admission, KV Cache residency, and cross-worker request orchestration. VenusRL's environment resource manager combines a memory-aware admission threshold with a template-keyed page-sharing pool, packing more sandboxes per node while preserving strict memory isolation via write-protected page table entry aliasing and copy-on-write. Across representative agentic RL workloads, VenusRL achieves up to 4.24x end-to-end training speedup over state-of-the-art baselines and reduces environment cost by up to 89%.