VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of purely data-driven video prediction, which accumulates errors and generates physical hallucinations over long horizons, thereby compromising robotic planning. To overcome this, we propose a Real-Sim-Real prediction framework grounded in digital twins. Specifically, this work pioneers the use of vision-language models (VLMs) to reconstruct executable digital twin scenes. By leveraging Isaac Sim with multiple randomized physics configurations, the framework provides dynamic constraints and contextual references, while conditional video generation incorporates authentic interaction details to establish a closed-loop predictive planning pipeline. The proposed approach effectively mitigates physical hallucinations in video prediction and achieves unified real-sim domain alignment, substantially improving the success rate and reliability of robotic manipulation planning.
📝 Abstract
While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins. For a target manipulation task, a VLM reconstructs an executable digital twin from a real demonstration episode. To accommodate the ill-posed estimation of unobserved physical properties, Isaac Sim simulates multiple forward dynamic rollouts across randomized physical configurations under candidate action trajectories. Using these rollouts as in-context references, VPTwin harmonizes both domains, using simulation dynamics to enforce physical plausibility while capturing unmodeled contact interactions from real video. Furthermore, we establish a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution. Evaluations show substantial reductions in physical hallucinations during video prediction and marked improvements in manipulation planning performance.
Problem

Research questions and friction points this paper is trying to address.

video prediction
robotic manipulation planning
compounding errors
physical hallucinations
world model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Prediction
Digital Twin
Robotic Manipulation Planning
Vision-Language Model
Sim-to-Real
🔎 Similar Papers
No similar papers found.