🤖 AI Summary
This work addresses the challenges of high-dimensional control, decentralized decision-making, and scalability in whole-body coordination for multiple humanoid robots by proposing a general multi-agent reinforcement learning framework. The method learns a shared high-level policy to integrate pretrained skills, enabling decentralized whole-body coordination using only task-level rewards. Furthermore, it introduces a permutation-based data augmentation strategy, theoretically proven to preserve policy gradient directions in homogeneous Markov games, thereby significantly improving learning efficiency. Experiments conducted on two distinct humanoid embodiments validate the framework's cross-task generalization capability, demonstrating superior performance and efficient coordinated behaviors across diverse tasks.
📝 Abstract
Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.