🤖 AI Summary
This study addresses the generalization challenges in humanoid robot control arising from strongly coupled action spaces and scarce real-world data by proposing VioLA. This method pioneers the use of motion latents rather than joint commands as the action space, executed via a pretrained controller to automatically annotate large-scale human videos into the robotic control domain, thereby eliminating reliance on conventional teleoperation data. Furthermore, VioLA integrates a unified motion encoder with vision-language-action (VLA) and world model backbones to enable direct learning of generalizable whole-body control from human videos. Experiments demonstrate that VioLA achieves zero-shot success rates of 100% for locomotion and 88.6% for manipulation tasks on a real humanoid robot, significantly outperforming baseline models such as GR00T N1.7.
📝 Abstract
Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.