EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出EBRL系统,通过异步管道调度器和细粒度资源管理技术解决强化学习中硬件资源利用效率低的问题。
📝 Abstract
Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training for efficiency, but exclusive GPU allocation and synchronized barrier in rollout still leave substantial hardware resource waste. In this paper, we present EBRL, an asynchronous embodied RL training system with two core techniques. The asynchronous pipelined scheduler overlaps rollout and training, pipelines simulation and generation across environment groups, and carries out each environment independently, eliminating synchronization stalls. The fine-grained resource manager pools CPU cores and GPU streaming multiprocessors, and uses stage profiles and runtime feedback to adjust resource quotas and batch sizes to meet the shifting demands among stages. We implement EBRL on RLinf and evaluate it with four embodied policies and four simulation benchmarks across heterogeneous GPU testbeds. Experiments show that EBRL achieves 1.30-3.47 times the end-to-end rollout throughput and 2.5 times of training convergency compared to the SOTA embodied RL systems.
Problem

Research questions and friction points this paper is trying to address.

Embodied Reinforcement Learning
Resource Management
Heterogeneous Demands
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

asynchronous pipelined scheduler
fine-grained resource manager
embodied reinforcement learning
🔎 Similar Papers
2024-07-09IEEE/ASME transactions on mechatronicsCitations: 94