Score
Design and implement a pipelined training infrastructure that groups environment rollouts into complete groups and immediately dispatches those groups to rollout workers so rollout generation overlaps with model training; the pipeline is scheduled to minimize trainer GPU idle time by continuously feeding training with freshly completed groups. Ensure the dispatch and buffering logic preserves on‑policy correctness by treating complete groups as atomic units for experience collection and update ordering.
This work addresses the low GPU utilization in existing synchronous on-policy reinforcement learning systems, which must wait for complete rollouts before training can commence, while asynchronous approaches improve hardware efficiency at the cost of policy degradation due to stale data. To reconcile efficiency and correctness, the authors propose RolloutPipe, a framework that decouples rollout and training through a Complete Group Pipeline (CGP) architecture with fixed rollout weights and a Front-Group Dispatching (FGD) mechanism. This design enables immediately dispatching trainable groups to the trainer upon generation, achieving efficient overlap between rollout and training while strictly preserving on-policy semantics. Experiments on the Qwen3-1.7B model demonstrate that RolloutPipe reduces end-to-end training time by 30.7%–42.3% and decreases trainer idle time by 37%–76% compared to the Slime system.
In reinforcement learning (RL) post-training, decoupling rollout and training improves hardware specialization but introduces severe inter-cluster idle time (“bubbles”) due to on-policy synchronization. Method: We propose RollMux, a cross-cluster cooperative scheduling framework featuring a novel *co-execution group* abstraction and a two-layer scheduler enabling phase-level multiplexing. It enforces *residency constraints* to keep model states persistently resident in memory, enabling low-overhead “warm-start” context switching. RollMux integrates conservative stochastic planning, provably optimal polling-based scheduling, locality-domain isolation, and GPU-cluster-wide coordinated orchestration. Results: Evaluated on a production-scale platform with 328×H20 and 328×H800 GPUs, RollMux achieves 1.84× higher cost efficiency than standard decoupled execution and 1.38× over the state-of-the-art co-location baseline, while meeting SLOs at 100% attainment.
This paper addresses two key challenges in industrial-scale reinforcement learning (RL): tight coupling between training and execution, and low GPU utilization with poor system scalability. To tackle these, we propose a decoupled, serverless RL framework. Our method separates the trainer and agent execution pipelines, introduces a data plane and trajectory management mechanism to enable seamless pausing and resuming of rollouts; designs a label-driven scheduler and spatiotemporal multiplexing pipeline to eliminate pipeline bubbles and unify heterogeneous resource scheduling; and dynamically reassigns idle training nodes to rollout tasks. Experimental results demonstrate that our framework significantly improves GPU utilization, reduces system idle time, and achieves high throughput, strong stability, and excellent scalability—particularly in complex scenarios such as multi-agent and long-horizon RL tasks.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.
本文提出EBRL系统,通过异步管道调度器和细粒度资源管理技术解决强化学习中硬件资源利用效率低的问题。
Miles v0.1系统旨在通过高效、可靠和可扩展的方式解决大规模强化学习的生产应用问题,采用SGLang构建引擎、双后端训练器及多种权重同步传输方法。
This study addresses the inherent trade-off in agent reinforcement learning between training engine idleness under synchronous execution and policy staleness under asynchronous execution. To reconcile this dilemma, we propose a fine-grained gradient flow mechanism that guarantees zero policy staleness. By processing single-trajectory or single-round data on the fly within Group Relative Policy Optimization (GRPO) and Online Policy Distillation (OPD), our method is theoretically proven to yield updates mathematically identical to those of batch synchronous training. Experimental results demonstrate that the proposed mechanism accelerates training by up to 1.9× compared to synchronous baselines while improving performance by up to 2.47 percentage points over asynchronous methods under a fixed computational budget, thereby successfully unifying training efficiency with optimization accuracy.
This work addresses the challenge of serving policy inference for multiple heterogeneous robots from a remote GPU, where conventional batched scheduling fails to accommodate disparities in action chunk consumption rates, thereby limiting system throughput. The authors formulate this scenario as a scheduling problem and introduce Armory, a system featuring the first batched scheduling algorithm that explicitly accounts for heterogeneity in action chunk consumption—departing from traditional homogeneous assumptions to better align with the realities of robotic closed-loop control. Armory integrates remote GPU batched inference, deployment of Vision-Language-Action models, and coordination mechanisms for heterogeneous robots. Evaluations on both real-world and simulated robot clusters demonstrate its efficacy, achieving up to an 18% improvement in overall throughput compared to naive scheduling strategies.
This study addresses the limitations of GPU memory capacity and poor scalability in multi-node large model training by proposing Scale-in, an architecture based on user-space transparent interception. The proposed method leverages CPU DRAM to extend GPU memory, achieving efficient data transfer through a zero-copy data path and RoCE network optimization. Furthermore, it seamlessly integrates with distributed parallel strategies, enabling training without modifying the underlying system. Experimental results demonstrate that Scale-in improves throughput by 42%–68% over state-of-the-art methods. Notably, it maintains 90% training efficiency while reducing GPU usage by 50% and achieves a 35% speedup in communication-bound scenarios, significantly lowering hardware requirements and network overhead.
论文提出了一种保真度感知的训练框架和C-DPPO方法,解决了编码和终端代理后训练中的令牌与控制保真度错误问题。