Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
This work addresses the challenge of learning robust and transferable world model representations from robotic visual data in complex outdoor environments. The authors propose a novel approach that, for the first time, integrates deep geometric priors with isotropy-induced latent space regularization (SIGReg) and incorporates an over-parameterization strategy during training to enhance the Joint Embedding Predictive Architecture’s (JEPA) capacity to model real-world scene dynamics while preserving inference efficiency. The method substantially outperforms the LeWM baseline, reducing visual odometry error by 33%, improving in-domain and out-of-domain anomaly detection separation on TartanGround, achieving higher fidelity in multi-step latent state prediction under domain shift, and demonstrating superior understanding of non-geometric physical factors such as illumination changes.