VioLA: Learning Generalist Humanoid Control Policies from Human Data

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization challenges in humanoid robot control arising from strongly coupled action spaces and scarce real-world data by proposing VioLA. This method pioneers the use of motion latents rather than joint commands as the action space, executed via a pretrained controller to automatically annotate large-scale human videos into the robotic control domain, thereby eliminating reliance on conventional teleoperation data. Furthermore, VioLA integrates a unified motion encoder with vision-language-action (VLA) and world model backbones to enable direct learning of generalizable whole-body control from human videos. Experiments demonstrate that VioLA achieves zero-shot success rates of 100% for locomotion and 88.6% for manipulation tasks on a real humanoid robot, significantly outperforming baseline models such as GR00T N1.7.
📝 Abstract
Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.
Problem

Research questions and friction points this paper is trying to address.

Humanoid Control
Generalist Policy
Action Space Coupling
Data Scarcity
Zero-shot Instruction Following
Innovation

Methods, ideas, or system contributions that make the work stand out.

Humanoid Control
Motion Latents
Zero-shot Transfer
Human Demonstrations
Vision-Language-Action Model
🔎 Similar Papers
2024-05-28International Conference on Learning RepresentationsCitations: 10
2024-07-16Neural Information Processing SystemsCitations: 16