RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high data cost and reliance on strong supervision in acquiring generalizable manipulation skills for humanoid robots operating in human-centric environments. The authors propose a skill distillation framework that requires neither human demonstrations nor teleoperation. Starting from monocular RGB-D videos, the method leverages a generative video model to synthesize human manipulation sequences, extracts keyframes via depth-aware 3D reconstruction, and geometrically preserves action retargeting onto high-DoF humanoid robots. A vision-language-model-guided optimization loop and online object-centric relocalization are further integrated to enable whole-body coordinated control. Experiments demonstrate that the system achieves cross-object generalization and robust recovery under perturbations on real robots, significantly enhancing the scalability and practicality of skill acquisition.
📝 Abstract
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.
Problem

Research questions and friction points this paper is trying to address.

humanoid manipulation
skill generalization
egocentric video
whole-body control
dexterous interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic skill distillation
egocentric video generation
whole-body manipulation
vision-language refinement
geometry-preserving retargeting