The Planning Limits of Latent World Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear effective prediction horizon and failure boundaries of latent world models in robotic planning. By constructing action-conditioned predictors atop five frozen visual backbones, including V-JEPA, and integrating the Meta-World simulation platform with model predictive control, it evaluates long-horizon planning capabilities across simulated and real-world environments. This work provides the first quantitative characterization of world model planning limits, revealing that long-horizon bottlenecks persist even under perfect predictions. Experiments demonstrate that pure imagination-based planning achieves only a 23% success rate, which is elevated to 76% through a proximal expert subgoal decomposition strategy. Furthermore, optimizing vision-language-action policy selection increases task success from 65% to 77%, offering practical insights for enhancing the reliability of world-model-based planners.
📝 Abstract
World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this question using action-conditioned predictors built on five frozen self-supervised visual backbones: V-JEPA 2, V-JEPA 2.1, VideoMAEv2, VideoPrism, and DINOv2. We use frozen backbones to test representations intended to transfer across environments. We evaluate these models on diverse Meta-World manipulation tasks and real-robot interactions from BridgeData V2. We find that a world model guides action selection reliably only when the goal lies within, or slightly beyond, the trajectory it imagines during planning. With five-step rollouts, the length the predictor was trained on, the world model ranks actions reliably only for targets five to ten control steps ahead, whereas task goals lie 16 to 53 steps away. Neither an 81-fold larger predictor nor longer-rollout training extends this range; the encoder affects both range and closed-loop success, with V-JEPA 2.1 performing most consistently. More fundamentally, the limit persists under perfect prediction: using the real simulator, success falls from 92% to 41% as the target moves from five to twenty steps ahead of a five-step rollout. Planning therefore requires either longer imagined trajectories or closer subgoals. For distant goals, pure imagination succeeds in 23% of episodes, planning with feedback (MPC) raises success to 30%, imagining as far as the goal to 47%, and nearby expert subgoals to 76%. Used within its plannable range, a world model can also improve a vision-language-action (VLA) policy: choosing among eight actions the VLA proposes raises its success from 65% to 77% across 16 different tasks.
Problem

Research questions and friction points this paper is trying to address.

world models
planning limits
robot manipulation
latent representations
trajectory prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent World Models
Planning Limits
Self-supervised Visual Backbones
Model Predictive Control
Vision-Language-Action Policy
🔎 Similar Papers
No similar papers found.