π€ AI Summary
Current JEPA-based world models rely on forward prediction and struggle to reliably disentangle a robotβs proprioceptive state from its dynamics, thereby limiting downstream planning performance. This work proposes the PSG-JEPA framework, which introduces a physics-guided mechanism during training: two complementary objectives anchor individual latent variables to the proprioceptive state and pairs of latent variables to joint-angle changes across multiple timescales, respectively. This design enhances the identifiability of physical states in the latent space without altering the inference pipeline. Experiments demonstrate that PSG-JEPA significantly outperforms existing world model baselines in latent interpretability, goal-conditioned planning with a frozen latent space, and policy learning both in simulation and on real robots.
π Abstract
Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.