Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of large-scale, diverse, and morphology-agnostic pretraining data for vision-language-action models. To bridge this gap, the authors propose a scalable synthetic framework that efficiently transforms massive first-person human demonstration videos into multimodal robot-centric data through motion retargeting, robotic arm visual synthesis, and multistage quality filtering. The resulting ego-to-robot dataset encompasses 15 distinct robot morphologies and totals 18,561 hours, accommodating both real-world and web-sourced video inputs. Integrated with a joint vision-language-action pretraining strategy and a decoupled perturbation evaluation framework, RoboTwin2.0, the approach significantly enhances out-of-distribution generalization across visual appearance, scene layout, embodiment morphology, and task semantics. Real-robot experiments validate the method’s effectiveness and practical transferability.
📝 Abstract
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present \textbf{Ego2Robot}, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/
Problem

Research questions and friction points this paper is trying to address.

robot manipulation
egocentric video
data synthesis
generalization
vision-language-action models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ego2Robot
action retargeting
robot data synthesis
vision-language-action models
out-of-distribution generalization
🔎 Similar Papers
No similar papers found.