JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prediction failure of pretrained world models caused by environmental dynamics shifts during testing. To this end, we propose a continual self-supervised test-time adaptation method built upon a joint-embedding predictive architecture. Specifically, only the latent dynamics predictor is updated to accommodate novel dynamics, while the visual encoder and reward head remain frozen. Furthermore, a dense replay buffer mechanism is designed to enable persistent adaptation without requiring target images or online rewards. Experimental results demonstrate that the proposed approach reduces prediction error by 83% on average and improves planning performance by 153%, significantly outperforming existing frozen baselines.
📝 Abstract
World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.
Problem

Research questions and friction points this paper is trying to address.

world models
test-time training
dynamics shifts
planning
latent dynamics predictor
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Training
World Models
Joint-Embedding Predictive Architecture
Dense Replay
Dynamics Shifts