train jepa world models

Design and train JEPA-style world models by building encoder networks that map raw inputs into compact latent scene embeddings, context/predictor modules that forecast future embedding trajectories, and appropriate contrastive or predictive losses to align predicted and target embeddings. Implement training pipelines, multisensor fusion and evaluation procedures that yield representations usable for unsupervised surprise/novelty scoring and other downstream analyses without labeled supervision.

trainjepaworldmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional generative modeling, which often focuses on pixel-level reconstruction and struggles to capture high-level semantics. To overcome this, the authors propose an Energy-Based Joint Embedding Predictive Architecture (EB-JEPA) that performs self-supervised prediction in representation space rather than pixel space, effectively enabling the construction of world models for images, videos, and action-conditioned environments. The study introduces the first open-source, lightweight, and modular EB-JEPA library, systematically demonstrating the critical role of regularization in preventing representational collapse. The framework supports multi-step temporal prediction and action-conditioned modeling, achieving strong empirical results: 91% probe accuracy on CIFAR-10, high-quality multi-step video prediction on Moving MNIST, and a 97% planning success rate on the Two Rooms navigation task—all trained within hours on a single GPU.

Energy-Based ModelsJoint-Embedding Predictive ArchitecturesRepresentation Learning

This work addresses the lack of theoretical understanding regarding the generalization capabilities of JEPA-based world models. By formulating JEPA pretraining as a conditional spectral graph learning problem, we characterize its learning objective through low-rank decomposition of an action-conditioned co-occurrence matrix and establish, for the first time, a theoretical link between pretraining error and downstream planning regret. We derive finite-sample generalization bounds for JEPA world models, revealing an intrinsic trade-off governed by the latent dimension between approximation error and sampling error. Our analysis provides the first theoretical foundation elucidating both the advantages and limitations of latent-space predictive models.

generalization theoryJEPAlatent predictive models

This work addresses the challenge of jointly modeling photometric invariance in images and temporal dynamics in videos within a unified framework. The authors propose UniJEPA, the first architecture that learns both image-level photometric prediction and video-level temporal state prediction end-to-end in a shared latent space, without relying on exponential moving averages (EMA), stop-gradient operations, or pretrained encoders. By combining next-embedding prediction loss with Gaussian regularization, UniJEPA achieves controllable abstraction: its photometric branch captures structural invariance, while its temporal branch learns dynamic equivariance. Experiments demonstrate that UniJEPA matches or exceeds the performance of specialized models across image, video, and control tasks, using only a single loss hyperparameter. Moreover, it enables zero-shot planning that is tens of times faster than generative world models while maintaining comparable accuracy.

Joint-Embedding Predictive Architecturelatent spaceself-supervised learning

This work addresses the challenge of predicting entirely missing objects in 3D scenes by introducing SR-JEPA, a scene-level point cloud joint-embedding predictive architecture. The method learns context-aware, queryable 3D latent states without requiring reconstruction, semantic labels, or 2D features, leveraging shape-agnostic, learnable query points anchored at object centers. As the first approach to achieve compositional predictive representations for fully absent objects, SR-JEPA employs a point-native JEPA framework with an exponential moving average target and a frozen prediction pathway. Experiments demonstrate its effectiveness, achieving a 43.13% semantic identity macro accuracy on ARKitScenes—surpassing the strongest baseline by 22.18 points—and attaining 41.15 AP on Sr3D, thereby validating the representational power and generalization capability of the learned latent states.

3D scenesjoint-embeddingmissing object

Graph-level Representation Learning with Joint-Embedding Predictive Architectures

Sep 27, 2023
GS
Geri Skenderi
🏛️ Bocconi University | Michigan State University | University of Verona

To address the limitations of contrastive learning—such as susceptibility to overfitting and difficulty in capturing semantic hierarchies—in graph-level representation learning, this paper proposes Graph-JEPA, the first adaptation of the Joint Embedding Predictive Architecture (JEPA) to graph-structured data. Graph-JEPA enables contrastive-free and reconstruction-free self-supervision by masking subgraphs and predicting their latent representations. Crucially, it introduces hyperbolic coordinate regression as a novel objective to explicitly model the implicit hierarchical structure among graph concepts. By eliminating negative sampling and pixel-level reconstruction, Graph-JEPA significantly mitigates overfitting. Extensive experiments demonstrate that Graph-JEPA consistently outperforms state-of-the-art self-supervised methods on graph classification, continuous-value regression, and non-isomorphic graph discrimination tasks. The learned graph-level representations exhibit superior semantic richness and generalization capability.

Graph FeaturesMachine LearningSelf-supervised Learning

Latest Papers

What's happening recently
View more

Existing general-purpose robotic policies model only short-term physical futures and lack explicit representations of stage-level semantic futures, hindering efficient planning for multi-stage manipulation tasks. This work proposes JEPA-WAM, which integrates a Stage-JEPA module into the Motus architecture to jointly model both short-term physical dynamics and stage-level semantic futures for the first time. The approach leverages a frozen V-JEPA2 encoder to extract stage-state representations and employs goal-conditioned joint embedding to predict latent representations of subsequent stages, thereby enhancing task-level planning capabilities. Evaluated on 50 tasks in RoboTwin 2.0, the method achieves an overall success rate of 90.25% and reduces action steps by an average of 5.97% among successful executions.

future predictiongeneralist robot policiesrobot manipulation

This work systematically identifies three distinct modes of representation collapse in JEPA-based world models—physical invariance, identifiability, and counterfactual dynamics—even when global latent collapse is avoided. To address these failure modes, the paper introduces PhyLatent, a novel training objective that jointly optimizes dynamics-relevant representations through physical state anchoring, future representation alignment, static visual invariance constraints, counterfactual branch disentanglement, and latent denoising. Moving beyond reliance on global non-collapse assumptions alone, PhyLatent significantly reduces the three collapse rates to 7.53%, 0.95%, and 4.62% on OGBench-Cube, yielding a model-predictive control (MPC) success rate of 78.1%. It further achieves a 98.0% success rate on the TwoRooms task and maintains state-of-the-art performance on Reacher and PushT benchmarks.

action consequencesdynamics-relevant representationsJEPA world models

This work addresses the challenge of learning robust and transferable world model representations from robotic visual data in complex outdoor environments. The authors propose a novel approach that, for the first time, integrates deep geometric priors with isotropy-induced latent space regularization (SIGReg) and incorporates an over-parameterization strategy during training to enhance the Joint Embedding Predictive Architecture’s (JEPA) capacity to model real-world scene dynamics while preserving inference efficiency. The method substantially outperforms the LeWM baseline, reducing visual odometry error by 33%, improving in-domain and out-of-domain anomaly detection separation on TartanGround, achieving higher fidelity in multi-step latent state prediction under domain shift, and demonstrating superior understanding of non-geometric physical factors such as illumination changes.

JEPAoutdoor environmentsreal-world data