๐ค AI Summary
This work addresses the challenge of learning robust and transferable world model representations from robotic visual data in complex outdoor environments. The authors propose a novel approach that, for the first time, integrates deep geometric priors with isotropy-induced latent space regularization (SIGReg) and incorporates an over-parameterization strategy during training to enhance the Joint Embedding Predictive Architectureโs (JEPA) capacity to model real-world scene dynamics while preserving inference efficiency. The method substantially outperforms the LeWM baseline, reducing visual odometry error by 33%, improving in-domain and out-of-domain anomaly detection separation on TartanGround, achieving higher fidelity in multi-step latent state prediction under domain shift, and demonstrating superior understanding of non-geometric physical factors such as illumination changes.
๐ Abstract
World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual complexity. This combines depth supervision with an isotropy-inducing latent regularizer (SIGReg), maximizing task-agnostic latent diversity while constraining how that diversity is organized, with the combined objective targeting the highest-entropy representation consistent with scene geometry. To satisfy this greater complexity without increasing inference time, we also add training-only overparameterization. Training an 18M-parameter model on video from a real agricultural robot, we evaluate with frozen-representation visual odometry probes, predictor-based surprise detection, and multi-step latent rollout fidelity. Compared to the baseline LeWM, our method lowers visual odometry probe error by 33%, substantially increases surprise-score separation both in-domain and on the out-of-domain TartanGround benchmark, and improves multi-step rollout fidelity under domain shift, with gains that grow with rollout horizon. Notably, we also see improvements in surprise-score separation on physics understanding that is not directly tied to 3D geometry, such as lighting and shadows. These results show that a lightweight training-time geometric prior makes a compact JEPA world model more useful and more transferable on real outdoor data with strong underlying representations, without adding inference overhead. Our work suggests that depth as a physically grounded prior can enhance world model generalization on a variety of tasks.