V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a novel paradigm for constructing efficient World Action Models (WAMs) based solely on predictive latent spaces, addressing the heavy reliance of existing WAMs on large generative visual models. Specifically, visual features are extracted using a frozen V-JEPA encoder, while a latent state predictor and a flow-matching action expert are trained from scratch. A key-value state conditioning mechanism is introduced to couple instruction-driven visual prediction with action generation. This study provides the first evidence that predictive latent features alone can support WAM learning without fully pretrained generative models. Despite having only 0.9B parameters, the proposed model achieves strong performance and robustness on benchmarks such as LIBERO. Furthermore, integrating unlabeled video pretraining significantly enhances its generalization capabilities.
📝 Abstract
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Predictive Visual Latents
Visual Foundation
Robot Manipulation
Out-of-Distribution Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Models
Predictive Visual Latents
V-JEPA
Flow-Matching Action Expert
Out-of-Distribution Generalization
🔎 Similar Papers
No similar papers found.