Frozen in a Frame: The Velocity Blind Spot in JEPA World Models

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of Joint Embedding Predictive Architecture (JEPA) world models, where single-frame inputs lack velocity information, thereby constraining planning performance. To overcome this “velocity blind spot,” we propose the TI-JEPA architecture, which decouples latent representations into pose and motion encodings for joint prediction. Additionally, we introduce RateIdent, a diagnostic protocol that explicitly models velocity dynamics without requiring privileged supervision. By leveraging finite-difference motion encoding and linear probe evaluation, our approach significantly reduces planning errors—by up to 55%—on benchmarks such as Pendulum. The proposed method substantially outperforms baselines while preserving interpretability.
📝 Abstract
Joint-embedding predictive architectures (JEPAs) for world modeling train an encoder so a predictor maps a current embedding and action to the next frame's embedding, always from a single rendered frame. This has a structural blind spot: a renderer without motion blur draws a scene from configuration alone, so a single-frame embedding carries no velocity information, for any encoder, including the official released LeWM weights. We confirm this on official checkpoints across four real benchmarks (PushT, Reacher, Cube, TwoRoom): every linear velocity probe sits at or below chance while position probes reach R^2 about 0.95. We introduce RateIdent, a three-stage diagnostic protocol, and TI-JEPA, a lightweight fix splitting the latent into a pose code and an explicit finite-difference motion code, predicted jointly. Across three physically grounded environments, TI-JEPA gives a significant, seed-robust gain on a stop-at-goal planning task over a matched-memory baseline, e.g. 55% lower final distance on Pendulum (p=3.2x10^-10) and 64% on CartPole (p=5.1x10^-15). We reproduce this at official ViT-Tiny plus AdaLN-transformer scale, then push the same recipe onto real dm_control Reacher photographs trained from scratch, where TI-JEPA's branch separation exceeds the memory-having baseline's by roughly 38x, the paper's largest margin. Against a same-footprint recurrent RSSM-style predictor, TI-JEPA matches or beats its rollout accuracy on two of three environments, stays separately probeable for pose and motion, and wins outright on the most coupled one. A checkable formal argument and six evaluated environments show single-frame targets are the wrong object to predict when velocity matters, and a small, interpretable structural change fixes it with no privileged supervision. Code, checkpoints, and the project page are linked below the title.
Problem

Research questions and friction points this paper is trying to address.

JEPA
world models
velocity blind spot
single-frame embedding
motion information
Innovation

Methods, ideas, or system contributions that make the work stand out.

JEPA
World Models
Velocity Blind Spot
Latent Disentanglement
TI-JEPA
🔎 Similar Papers