DeepJEPA: Scaling World Models from Within

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational inefficiency of existing world model planners, which uniformly allocate resources regardless of decision complexity. We propose DeepJEPA, a framework that reconceptualizes test-time scaling as an internal compute allocation problem by treating transition depth as the scaling axis. Built upon a weight-shared Joint Embedding Predictive Architecture (JEPA) integrated with an adaptive recurrence mechanism, DeepJEPA dynamically performs recursive updates exclusively at decision-critical junctures to optimize computational allocation. Evaluated across five visual control tasks, our approach requires only 1.00–1.26 updates on average to match or surpass the performance of fixed-depth planners. These results demonstrate that DeepJEPA achieves efficient and precise test-time compute scaling, offering a principled alternative to uniform computation strategies in model-based planning.
📝 Abstract
World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only 1.00-1.26 updates per transition. Its additional computation concentrates at contact onset and sustained object interaction, where latent corrections can change which candidates enter the planner's elite set and which action is selected. Representation probes further show that improved planning does not require uniformly better object-state decodability. DeepJEPA therefore reframes world-model scaling as a problem of allocating internal computation where it can change the planner's decision: think deeper at decision-critical transitions instead of making every rollout uniformly deeper or longer.
Problem

Research questions and friction points this paper is trying to address.

world models
test-time scaling
computation allocation
planning
decision-critical transitions
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Joint-Embedding Predictive Architecture
Test-Time Scaling
Adaptive Computation
Visual Control
🔎 Similar Papers
No similar papers found.