Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity and high acquisition cost of high-quality interaction data for humanoid mobile manipulation by proposing the PRISM framework. This method leverages video generation models to synthesize counterfactual videos for training data augmentation, establishing a "real-sim-real" closed loop. Specifically, it converts generated videos into simulated interaction data via contact-anchored reconstruction and motion retargeting, subsequently training control policies through reinforcement learning combined with deep visual perception. Experimental results demonstrate that, requiring only a few real-world demonstration videos, the proposed framework achieves zero-shot generalization across object categories and scene configurations, successfully accomplishing pick-and-carry tasks.
📝 Abstract
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Problem

Research questions and friction points this paper is trying to address.

humanoid loco-manipulation
visual imitation
interaction video collection
scalable robot learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Video Generation
Real-to-Sim-to-Real
Humanoid Loco-Manipulation
Video-to-Video Generation
Contact-Anchored Retargeting