Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data

๐Ÿ“… 2026-07-14
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of learning robust and transferable world model representations from robotic visual data in complex outdoor environments. The authors propose a novel approach that, for the first time, integrates deep geometric priors with isotropy-induced latent space regularization (SIGReg) and incorporates an over-parameterization strategy during training to enhance the Joint Embedding Predictive Architectureโ€™s (JEPA) capacity to model real-world scene dynamics while preserving inference efficiency. The method substantially outperforms the LeWM baseline, reducing visual odometry error by 33%, improving in-domain and out-of-domain anomaly detection separation on TartanGround, achieving higher fidelity in multi-step latent state prediction under domain shift, and demonstrating superior understanding of non-geometric physical factors such as illumination changes.
๐Ÿ“ Abstract
World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual complexity. This combines depth supervision with an isotropy-inducing latent regularizer (SIGReg), maximizing task-agnostic latent diversity while constraining how that diversity is organized, with the combined objective targeting the highest-entropy representation consistent with scene geometry. To satisfy this greater complexity without increasing inference time, we also add training-only overparameterization. Training an 18M-parameter model on video from a real agricultural robot, we evaluate with frozen-representation visual odometry probes, predictor-based surprise detection, and multi-step latent rollout fidelity. Compared to the baseline LeWM, our method lowers visual odometry probe error by 33%, substantially increases surprise-score separation both in-domain and on the out-of-domain TartanGround benchmark, and improves multi-step rollout fidelity under domain shift, with gains that grow with rollout horizon. Notably, we also see improvements in surprise-score separation on physics understanding that is not directly tied to 3D geometry, such as lighting and shadows. These results show that a lightweight training-time geometric prior makes a compact JEPA world model more useful and more transferable on real outdoor data with strong underlying representations, without adding inference overhead. Our work suggests that depth as a physically grounded prior can enhance world model generalization on a variety of tasks.
Problem

Research questions and friction points this paper is trying to address.

world models
JEPA
real-world data
visual complexity
outdoor environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

depth regularization
JEPA world models
geometric prior
latent regularizer
transferable representations
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
U
Usman M. Khan
Aigen