Score
Designs and implements hierarchical joint-embedding predictive architectures (JEPA) that learn multi-level representations by jointly encoding multiple abstraction levels and predicting embeddings across temporal scales; variants such as ER-JEPA / event reconstruction JEPA include modules to reconstruct or predict event-level embeddings. Uses self-supervised pretraining on unlabeled multivariate time series to produce representations that improve downstream task performance with limited labeled data.
Existing self-supervised time series representation learning methods—such as masked modeling—are vulnerable to input noise and confounding variables. To address this, we propose Time Series Joint Embedding Predictive Architecture (TS-JEPA), the first framework to adapt the JEPA paradigm to time series. TS-JEPA jointly optimizes future segment prediction and contrastive representation learning in a latent space, eliminating reliance on raw-input reconstruction and thereby significantly enhancing robustness. Its unified architecture natively supports both classification and forecasting tasks. Evaluated across multiple standard benchmarks, TS-JEPA achieves state-of-the-art or competitive performance, demonstrating strong generalization and balanced multi-task capability. This work establishes a novel paradigm for developing robust, general-purpose time series foundation models.
To address the limitations of contrastive learning—such as susceptibility to overfitting and difficulty in capturing semantic hierarchies—in graph-level representation learning, this paper proposes Graph-JEPA, the first adaptation of the Joint Embedding Predictive Architecture (JEPA) to graph-structured data. Graph-JEPA enables contrastive-free and reconstruction-free self-supervision by masking subgraphs and predicting their latent representations. Crucially, it introduces hyperbolic coordinate regression as a novel objective to explicitly model the implicit hierarchical structure among graph concepts. By eliminating negative sampling and pixel-level reconstruction, Graph-JEPA significantly mitigates overfitting. Extensive experiments demonstrate that Graph-JEPA consistently outperforms state-of-the-art self-supervised methods on graph classification, continuous-value regression, and non-isomorphic graph discrimination tasks. The learned graph-level representations exhibit superior semantic richness and generalization capability.
This work addresses the representation collapse problem in vision-based reinforcement learning (RL), caused by the entanglement of visual representation learning and policy optimization. We introduce the Joint-Embedding Predictive Architecture (JEPA)—a self-supervised framework—into RL for the first time, proposing a decoupled representation learning mechanism: a vision Transformer is employed to construct JEPA’s predictive objective, explicitly separating perceptual modeling from policy optimization to mitigate representation degradation. Evaluated on dynamic control benchmarks including CartPole, our approach significantly improves training stability. The robust, JEPA-derived visual embeddings serve as high-quality inputs for downstream policy learning, enabling end-to-end policies with superior performance and generalization. This work establishes a novel paradigm for self-supervised, representation-driven visual RL.
This work addresses the challenge of effectively leveraging multivariate time-series data—such as 12-lead electrocardiograms (ECGs)—in medical domains where labeled data are scarce. The authors propose the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), which introduces a novel Hierarchical JEPA (H-JEPA) model. This architecture employs a two-stage hierarchical design: it first models local temporal segments and then treats the resulting representations as a univariate sequence for global modeling, all within a Vision Transformer backbone trained via self-supervised learning. Inspired by clinical ECG interpretation workflows, the approach enables multi-level abstract representation learning. Pretrained on only approximately 180,000 ten-second ECG recordings, the model achieves state-of-the-art performance on the ST-MEM benchmark while maintaining computational efficiency and low resource consumption.
This work addresses the disconnect between Joint Embedding Predictive Architectures (JEPA) and probabilistic generative modeling, as well as JEPA’s reliance on heuristic regularization to prevent representation collapse. From a variational inference perspective, we reinterpret JEPA as a deterministic special case of a coupled latent variable model and, for the first time, integrate it into a variational autoencoding framework, yielding Var-JEPA. This approach introduces an explicit generative structure that unifies predictive and generative self-supervised learning, enabling meaningful representations without heuristic anti-collapse regularizers and supporting uncertainty quantification in the latent space. Leveraging ELBO optimization and a context–target prediction architecture, we instantiate Var-T-JEPA for tabular data, which significantly outperforms T-JEPA on downstream tasks and matches the performance of strong baseline methods using raw features.
This work addresses the challenge of jointly modeling photometric invariance in images and temporal dynamics in videos within a unified framework. The authors propose UniJEPA, the first architecture that learns both image-level photometric prediction and video-level temporal state prediction end-to-end in a shared latent space, without relying on exponential moving averages (EMA), stop-gradient operations, or pretrained encoders. By combining next-embedding prediction loss with Gaussian regularization, UniJEPA achieves controllable abstraction: its photometric branch captures structural invariance, while its temporal branch learns dynamic equivariance. Experiments demonstrate that UniJEPA matches or exceeds the performance of specialized models across image, video, and control tasks, using only a single loss hyperparameter. Moreover, it enables zero-shot planning that is tens of times faster than generative world models while maintaining comparable accuracy.
This work addresses the susceptibility of existing graph self-supervised learning methods to low-level input statistics and their limited capacity to model structural relationships among nodes. It introduces, for the first time, the Joint-Embedding Predictive Architecture (JEPA) paradigm to node-level graph representation learning through a structure-conditioned prediction mechanism: by masking k-hop subgraph structures, a context encoder predicts the target node’s representation in latent space, thereby circumventing reliance on input reconstruction or handcrafted augmentations. The approach integrates an EMA target encoder, cross-attention over spectral and centrality descriptors, and variance/covariance/Laplacian regularization, complemented by a progressive curriculum masking strategy to explicitly reinforce structural information learning. Evaluated on standard node classification benchmarks, the method achieves strong performance under both linear probing and fine-tuning, with ablation studies confirming the contribution of each component.