Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing dexterous manipulation datasets struggle to simultaneously achieve large scale and embodied alignment: teleoperated data is well-aligned but scarce, simulation is scalable yet suffers from domain gaps, and in-the-wild first-person human videos are abundant but mismatched with robotic deployment. To bridge this gap, the paper introduces a generative video world model that synthesizes large-scale, first-person human manipulation videos conditioned on language, objects, and scenes. It then converts these videos into robot-trainable supervision signals via hand motion reconstruction and visual editing, jointly fine-tuning vision-language-action (VLA) models with a small amount of real robot data. Evaluated on 18 real-world dexterous tasks, this approach significantly improves zero-shot success rates from 8.3% to 38.9%, effectively closing the divide between large-scale human video data and robot embodiment.
📝 Abstract
Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with robot deployment but expensive to collect; simulation is scalable but limited by the sim-to-real gap; and real egocentric videos scale effectively but remain misaligned with robot deployment. We propose Wh0, a framework that uses generative video world models as scalable and controllable sources of egocentric human-hand manipulation data to unlock the manipulation capabilities of pretrained dexterous VLA models. Conditioned on language, objects, and scenes, Wh0 uses a generative world model to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. Wh0 then converts the generated videos into robot-trainable supervision through hand motion reconstruction and visual editing. Co-trained with a limited amount of real robot data, WM-H adapts pretrained VLA models to dexterous manipulation deployment. Across 18 real-world dexterous manipulation tasks, compared with a model post-trained only on robot data, Wh0 improves zero-shot success on unseen tasks from 8.3% to 38.9%. Ablation studies further show that scalable generation and scene/embodiment alignment are key drivers of performance gains. Videos and open-source code can be found on our project website: https://chenyt31.github.io/wh0.github.io/.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
data scalability
embodiment alignment
egocentric video
sim-to-real gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative world models
egocentric human manipulation
dexterous manipulation
sim-to-real transfer
vision-language-action (VLA) models