Robotizing Human Videos with Physically Consistent Interactions

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively transferring human videos to robot policies, which is hindered by embodiment discrepancies between human and robot hands as well as interaction inconsistencies and occlusion artifacts introduced by video editing. To overcome these limitations, this work proposes a data generation framework based on contact reconstruction and depth-aware synthesis. Specifically, dense 3D contact prediction is employed to ensure stable grasping, while depth information preserves physically consistent occlusion relationships. By integrating hand-object segmentation with co-training via a diffusion policy visual encoder, the framework transforms human videos into high-quality robot manipulation data. Evaluations on the RoboTwin benchmark demonstrate that the proposed approach achieves state-of-the-art success rates and significantly enhances policy robustness against out-of-distribution scene variations and visual disturbances.
📝 Abstract
Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.
Problem

Research questions and friction points this paper is trying to address.

embodiment gap
human videos
robot manipulation
interaction reconstruction
physically consistent compositing
Innovation

Methods, ideas, or system contributions that make the work stand out.

embodiment gap
contact reconstruction
depth-aware compositing
co-training
robotizing human videos