VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training throughput bottlenecks caused by uneven trajectory completion in agent reinforcement learning, as well as resource waste stemming from static memory allocation in tool sandboxes. To overcome these challenges, we propose a fully decoupled architecture. Methodologically, we design a length-prediction-based, priority-aware action scheduler that accelerates the completion of critical sample groups to unlock training steps. Furthermore, we construct a template-keyed page-sharing pool that leverages copy-on-write and write-protected aliasing mechanisms to achieve efficient isolation and reuse of environment memory. Experimental results demonstrate that the proposed system achieves up to 4.24× end-to-end training speedup over state-of-the-art baselines while reducing environment resource overhead by 89%.
📝 Abstract
Agentic Reinforcement Learning (RL) trains LLM agents through multi-turn interactions with external tool environments. Its multi-turn nature exposes two system-level bottlenecks unaddressed by existing agentic RL frameworks. First, end-to-end training throughput is constrained by the slowest trajectories to complete, yet optimizing per-GPU utilization alone scatters rollout progress across many groups, delaying the completion of enough groups to unblock the next training step. Second, tool sandboxes are statically over-provisioned by their declared memory ceilings, leaving most physical memory stranded while replicating near-identical state across sandboxes launched from the same prompt. We present VenusRL, a fully disaggregated agentic RL system that addresses both bottlenecks. VenusRL's priority-aware action scheduler uses length-prediction heuristics to identify sample groups whose completion is most likely to unblock the next training step, and pushes them ahead of others across batch admission, KV Cache residency, and cross-worker request orchestration. VenusRL's environment resource manager combines a memory-aware admission threshold with a template-keyed page-sharing pool, packing more sandboxes per node while preserving strict memory isolation via write-protected page table entry aliasing and copy-on-write. Across representative agentic RL workloads, VenusRL achieves up to 4.24x end-to-end training speedup over state-of-the-art baselines and reduces environment cost by up to 89%.
Problem

Research questions and friction points this paper is trying to address.

Agentic Reinforcement Learning
Training Throughput
Tool Sandbox
Memory Over-provisioning
System Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
Disaggregated System
Priority Scheduling
Page Sharing
Copy-on-Write
🔎 Similar Papers
No similar papers found.
M
Mingjun Zhang
Institute of Computing Technology, Chinese Academy of Sciences
Yucheng Li
Yucheng Li
University of Surrey
Language Processing
M
Menghao Zhang
Beihang University
S
Shuyong Zhu
Institute of Computing Technology, Chinese Academy of Sciences
P
Ping Zhang
Infrawaves
Xiaohe Hu
Xiaohe Hu
Tsinghua University
machine learningsystem and architecture
J
Jun Chen
Infrawaves
Zhixin Wang
Zhixin Wang
ZheJiang University
RL systems
X
Xutong Wang
Infrawaves
H
He Liu
Infrawaves
Y
Yanmin Jia
Infrawaves
S
Shengrong Zhu
Infrawaves
P
Peng Sun
Shanghai Qiji Zhifeng Co., Ltd.
Mingjie Zhang
Mingjie Zhang
MPhil Student, The Hong Kong University of Science and Technology (Guangzhou)
RoboticsVision-Language Navigation
L
Liming Liu
Shanghai Innovation Institute
Jinlong Hou
Jinlong Hou
Shanghai Innovation Institute (SII)
machine learningdeep learninghigh performance computingdrug discoverymedical
Y
Yuan Cheng
Shanghai Innovation Institute
Y
Yujun Zhang
Institute of Computing Technology, Chinese Academy of Sciences