Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

πŸ“… 2026-08-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current JEPA-based world models rely on forward prediction and struggle to reliably disentangle a robot’s proprioceptive state from its dynamics, thereby limiting downstream planning performance. This work proposes the PSG-JEPA framework, which introduces a physics-guided mechanism during training: two complementary objectives anchor individual latent variables to the proprioceptive state and pairs of latent variables to joint-angle changes across multiple timescales, respectively. This design enhances the identifiability of physical states in the latent space without altering the inference pipeline. Experiments demonstrate that PSG-JEPA significantly outperforms existing world model baselines in latent interpretability, goal-conditioned planning with a frozen latent space, and policy learning both in simulation and on real robots.
πŸ“ Abstract
Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.
Problem

Research questions and friction points this paper is trying to address.

world models
latent representations
physical state grounding
JEPA
forward prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

JEPA
physical state grounding
latent identifiability
world models
robotic representation learning
πŸ”Ž Similar Papers
No similar papers found.
Haodong Yan
Haodong Yan
PhD student of INTR, HKUST (GZ)
Human reconstructionmotion prediction
J
Jiaguan Zhu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
M
Mingyuan Jia
COCO Matrix
R
Ruiqing Yin
COCO Matrix
Junjie He
Junjie He
Guizhou University
MRIDeep LearningCT
Zhide Zhong
Zhide Zhong
Beijing Institute of Technology
Robotics
J
Junfeng Li
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
J
Jinxuan Lu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
H
Hengtao Li
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
T
Tianran Zhang
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Jiayi Chen
Jiayi Chen
cuhksz
AIRoboticsControl
Wenxuan Song
Wenxuan Song
The Hong Kong University of Science and Technology (Guangzhou)
Vision-language-action ModelRobotics
Wen Chen
Wen Chen
PhD, The Chinese University of Hong Kong
Point Cloud RegistrationSLAMState Estimation
Yuxiang Gao
Yuxiang Gao
Johns Hopkins University
RoboticsHuman-Robot InteractionSocially-aware Navigation
Haoang Li
Haoang Li
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Robotics3D Computer Vision