Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of planning whole-body, contact-rich manipulation for humanoid robots using first-person RGB and proprioceptive inputs. The authors propose an end-to-end framework that employs a visual planner to predict torso, wrist, and foot poses, coupled with a reinforcement learning-based whole-body tracker for action execution. Notably, the approach enables direct training from human demonstrations without requiring motion retargeting and introduces a cross-error augmentation mechanism to enhance multi-module coordination robustness. Experiments on a real Unitree G1 robot demonstrate successful execution of complex tasks, including box-kicking sorting, object catching and throwing, and climbing. In simulation, the system achieves a 77% success rate for undisturbed sorting and maintains 64% under 40 Nยทs impulse disturbances, validating its effectiveness in dynamic, contact-intensive scenarios.
๐Ÿ“ Abstract
Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N.s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.
Problem

Research questions and friction points this paper is trying to address.

Humanoid robots
Whole-body manipulation
Egocentric vision
Proprioception
Contact-rich planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Whole-Body Loco-Manipulation
Human Demonstrations
Reinforcement Learning
Visual Planner
Cross-Policy Data Augmentation
๐Ÿ”Ž Similar Papers
No similar papers found.