🤖 AI Summary
This study addresses the performance degradation in planning caused by the entanglement of dynamics and background information within JEPA-based world models. We propose a lightweight, end-to-end world model that employs differentiable residual connections to route information into predictive latents for latent-space planning and residual context embeddings for reconstruction. By integrating self-supervised learning with cross-attention mechanisms, we construct an efficient discretized representation architecture and theoretically prove that this decoupled representation satisfies sufficiency, minimality, and invariance. Experiments demonstrate that our method improves planning success rates by 9% in simulated environments. Furthermore, in real-world robotic tasks, it surpasses substantially larger models while utilizing only 5.5M parameters, achieving significantly more efficient planning.
📝 Abstract
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.