Score
Design and implement GPU‑batched differentiable rollout and simulation pipelines that execute large numbers of parallel environment rollouts, compute efficient batched forward and backward passes across time, and scale to millions of environment steps. Use these pipelines to build gradient‑based training and analysis methods, including procedures for training under deployment perturbations and for robustness to timing or scheduling variations during rollouts.
In large-scale parallel GPU-based reinforcement learning, synchronous environment resets induce non-stationarity in the state distribution, leading to biased policy gradients and training instability. To address this, we propose Staggered Reset—a lightweight mechanism that introduces temporal diversity into rollout sequences by asynchronously initializing environments and randomizing reset timings, without incurring additional computational overhead. The method is fully compatible with standard on-policy algorithms such as PPO and requires no modifications to network architecture or loss functions. Empirical evaluation on high-dimensional robotic control tasks demonstrates that Staggered Reset improves sample efficiency, accelerates wall-clock convergence, and enhances final policy performance. Crucially, the gains scale positively with parallelism—larger numbers of concurrent environments yield greater improvements—thereby strengthening both stability and scalability of massively parallel RL training.
Reinforcement learning (RL) training of large language models (LLMs) on heterogeneous GPU clusters suffers from low resource utilization due to significant disparities in computational intensity, memory demand, and communication patterns across the three sequential phases—rollout generation, reward computation, and policy update. Method: This paper proposes AReaL-Hex, a system featuring a fully asynchronous RL architecture that decouples these three phases, coupled with a two-stage scheduler integrating mixed-integer linear programming (MILP)-based constraint search and graph partitioning to dynamically assign HBM/I/O-intensive and compute-intensive tasks to optimal heterogeneous devices while preserving data freshness. Contribution/Results: Experiments on 1.5B–14B LLMs demonstrate that AReaL-Hex achieves a 1.50× higher training throughput under identical budget constraints and reduces training cost by 1.46× at equivalent throughput, significantly improving efficiency on heterogeneous GPU clusters.
To address low resource utilization and limited throughput in video analytics inference services on heterogeneous GPU clusters, this paper proposes PPipe—a novel system that introduces pooled pipeline parallelism to latency-sensitive inference for the first time. Leveraging the observation that low-end and high-end GPUs exhibit comparable per-layer inference latency for multi-layer models, PPipe enables cross-device collaborative computation. It features an MILP-based optimization control plane, a resource reservation mechanism, and an adaptive dynamic batching strategy to unify scheduling across heterogeneous accelerators. Evaluation across 18 CNN models demonstrates that PPipe improves low-end GPU utilization by 41.1%–65.5%, increases end-to-end service throughput by 32.2%–75.1% over baseline systems, and significantly enhances cooperative efficiency and scalability of inference services on heterogeneous hardware.
To address the low parallel efficiency and suboptimal Tensor Core utilization of irregular sparse computations—such as Mixture-of-Experts (MoE)—on GPUs, this paper proposes a novel execution paradigm that synergistically combines static batching with dynamic task mapping. It statically compiles a dense task graph during compilation, transforming dynamic sparse inference into a single-kernel execution; a lightweight runtime scheduler then enables fine-grained task mapping onto hardware resources. This approach achieves, for the first time, highly efficient, targeted Tensor Core computation for MoE inference, attaining 91% and 95% of peak throughput utilization on NVIDIA H800 and H20 GPUs, respectively—significantly outperforming existing dynamic batching methods. The core contribution is a pioneering “compiler–runtime” co-optimization framework, establishing a new paradigm for high-throughput deployment of sparse models on hardware accelerators.
Frequent fine-grained kernel launches on GPUs incur substantial launch overhead, severely limiting performance in scientific computing. To address this, we propose a synergistic optimization combining iterative batching and CUDA Graph unrolling: multiple iterations are grouped into batches and statically unrolled into a single CUDA Graph, thereby eliminating redundant kernel launch overhead. We further introduce the first platform-agnostic criterion for selecting the optimal batch size and develop a generalizable analytical performance model. Evaluated on skeleton applications, our approach achieves over 1.4× speedup. It demonstrates significant and robust performance improvements across real-world iterative GPU applications—including Hotspot, Hotspot3D, and an FDTD-based Maxwell solver—without requiring application-specific tuning. This work establishes a general, analytically tractable, low-overhead execution paradigm for iterative GPU computations.
This work addresses the limitations of current robotic reinforcement learning systems, which rely heavily on GPU-based centralized simulation constrained by the CUDA ecosystem. The authors propose UniLab, a heterogeneous architecture that achieves the first efficient decoupling of CPU-parallelized simulation and GPU-accelerated policy learning. By introducing a unified runtime to manage data movement, buffering, and synchronization, UniLab establishes an end-to-end training loop across diverse hardware platforms. The framework supports non-CUDA environments—including macOS, ROCm, and Intel XPU—and integrates CPU-batched physics backends (MuJoCoUni and MotrixSim) alongside mainstream RL algorithms such as PPO, SAC, and TD3. Experimental results demonstrate a 3–10× improvement in training efficiency over existing approaches under identical hardware conditions, substantially overcoming prevailing platform and performance bottlenecks.
This work addresses the challenges of low communication efficiency and system complexity in large-scale AI training across geographically distributed data centers. It presents the first systematic characterization and joint optimization of three critical dimensions in “scale-across” training: parallelism strategy deployment, job scheduling, and network transport. By co-designing parallel placement, scheduling policies, and advanced networking techniques—and validating the approach through both real-world testbeds and large-scale simulations—the study achieves end-to-end global optimization of computation and communication. Experimental results demonstrate that the proposed solution improves training throughput by up to 64.62% over current production configurations and by 37.59% compared to state-of-the-art baselines.
This work addresses the low GPU utilization in existing synchronous on-policy reinforcement learning systems, which must wait for complete rollouts before training can commence, while asynchronous approaches improve hardware efficiency at the cost of policy degradation due to stale data. To reconcile efficiency and correctness, the authors propose RolloutPipe, a framework that decouples rollout and training through a Complete Group Pipeline (CGP) architecture with fixed rollout weights and a Front-Group Dispatching (FGD) mechanism. This design enables immediately dispatching trainable groups to the trainer upon generation, achieving efficient overlap between rollout and training while strictly preserving on-policy semantics. Experiments on the Qwen3-1.7B model demonstrate that RolloutPipe reduces end-to-end training time by 30.7%–42.3% and decreases trainer idle time by 37%–76% compared to the Slime system.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.
This work addresses the limitations of existing static batching strategies in handling bursty and heterogeneous inference requests, which struggle to dynamically balance throughput and latency. The authors formulate inference batching and routing as a Markov decision process and train reinforcement learning agents—using REINFORCE and PPO algorithms—on a discrete-event simulator grounded in queueing theory and real-world production traces. The agents make dynamic scheduling decisions based on queue states, request types, and GPU availability. Experimental results demonstrate that, in heterogeneous multi-GPU settings, the proposed approach improves throughput by up to 60% and reduces latency by 25% compared to heuristic policies such as round-robin and shortest queue. Under SLA compliance constraints, performance gains reach as high as 348%. Moreover, the agent autonomously learns an effective workload isolation policy that mitigates head-of-line blocking, highlighting the advantages of reinforcement learning in orchestrating complex, multi-resource scheduling scenarios.