Video Prediction Policy 2: Predict Better, Act Better

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited video prediction accuracy of existing world action models in open environments, which leads to erroneous robotic execution. We propose VPP2-14B, a video prediction policy model that enhances representations through continued pre-training on large-scale manipulation videos and event-level annotations. A Mixture-of-Transformers (MoT) architecture is designed to effectively decouple video generation from action control by integrating single-step visual planner distillation with an implicit inverse dynamics module. The resulting model exhibits strong zero-shot generalization capabilities, improving instruction-following rates by 11.0% and surpassing baseline performance by 18.5% in success rate on real-world ALOHA tasks, while achieving leading results across multiple benchmarks.
📝 Abstract
World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textit{event-level} video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0\% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5\% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Video Prediction
Robot Manipulation
Zero-shot Generalization
Motion Prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Event-level Pretraining
Mixture-of-Transformers
Zero-shot Generalization
Visual Planner
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yanjiang Guo
Yanjiang Guo
Tsinghua University
Embodied AIGenerative Model
Haodong Yan
Haodong Yan
PhD student of INTR, HKUST (GZ)
Human reconstructionmotion prediction
Zhide Zhong
Zhide Zhong
Beijing Institute of Technology
Robotics
Z
Zhongru Zhang
Tsinghua University
Q
Qingyuan Yang
Robotera, Tsinghua University
Q
Qingzhou Lu
Robotera, Tsinghua University
Xiaoyu Chen
Xiaoyu Chen
Tsinghua University
Embodied AIReinforcement Learning
Yen-Jen Wang
Yen-Jen Wang
UC Berkeley
Robotics
S
Shuying Deng
Robotera, Tsinghua University, University of California, Berkeley
C
Chenghan Yang
Tsinghua University
P
Puzhen Yuan
Robotera, Tsinghua University
C
Chenxin Liu
Robotera, Tsinghua University
T
Tun Ban
Robotera, Shanghai Jiaotong University
Xiang Zhu
Xiang Zhu
Institute for Interdisciplinary Information Sciences, Tsinghua University
robotics
Y
Yichen Liu
Robotera, Tsinghua University
Kun Feng
Kun Feng
Illinois Institute of Technology
Haoang Li
Haoang Li
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Robotics3D Computer Vision
Jianyu Chen
Jianyu Chen
Assistant Professor, Tsinghua University
AIRobotics