🤖 AI Summary
This study addresses the challenges of low sample efficiency and gradient conflicts in multi-task learning for quadruped robots by proposing a three-stage training framework. First, task-specific teacher policies are trained independently. Second, an adversarial task selection mechanism is introduced to dynamically prioritize the worst-performing tasks, thereby mitigating optimization conflicts. Finally, a unified student policy is generated through multi-teacher knowledge distillation that combines reinforcement learning with imitation learning. The proposed approach enables a single policy to master 22 distinct locomotion skills with significantly higher quality than baseline methods. Furthermore, the framework is successfully deployed on a real-world Unitree B1 robot, demonstrating its effectiveness and generalization capability in complex multi-task scenarios.
📝 Abstract
Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional stages train a student policy with a multi-teacher distillation setup that uses a combined RL and Imitation Learning (IL) objective under an adversarial task selection process that focuses training on the worst-performing task. With this method, we train a student policy that performs 22 tasks using 8 teachers. Evaluations show our method preserves motion quality and tracks commands more accurately than PPO and distill-then-finetune baselines, and in some cases generalizes to new tasks without explicit training. Finally, we demonstrate real-world robustness by deploying the resulting policy on a Unitree B1 quadruped. Video: https://youtu.be/V9yX04EBcFA