jepa pretraining

Designs and implements self-supervised pretraining pipelines based on joint-embedding predictive architectures (JEPA) that train an encoder to predict target latent embeddings from another augmented or masked view of the same input. This involves choosing view/augmentation strategies, encoder and target networks, predictive losses and regularization (e.g., target-network or embedding-space regularizers) to produce transferable, fine-grained latent representations that preserve relevant structural correlations for downstream tasks.

jepapretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Graph-level Representation Learning with Joint-Embedding Predictive Architectures

Sep 27, 2023
GS
Geri Skenderi
🏛️ Bocconi University | Michigan State University | University of Verona

To address the limitations of contrastive learning—such as susceptibility to overfitting and difficulty in capturing semantic hierarchies—in graph-level representation learning, this paper proposes Graph-JEPA, the first adaptation of the Joint Embedding Predictive Architecture (JEPA) to graph-structured data. Graph-JEPA enables contrastive-free and reconstruction-free self-supervision by masking subgraphs and predicting their latent representations. Crucially, it introduces hyperbolic coordinate regression as a novel objective to explicitly model the implicit hierarchical structure among graph concepts. By eliminating negative sampling and pixel-level reconstruction, Graph-JEPA significantly mitigates overfitting. Extensive experiments demonstrate that Graph-JEPA consistently outperforms state-of-the-art self-supervised methods on graph classification, continuous-value regression, and non-isomorphic graph discrimination tasks. The learned graph-level representations exhibit superior semantic richness and generalization capability.

Graph FeaturesMachine LearningSelf-supervised Learning

This work addresses the disconnect between Joint Embedding Predictive Architectures (JEPA) and probabilistic generative modeling, as well as JEPA’s reliance on heuristic regularization to prevent representation collapse. From a variational inference perspective, we reinterpret JEPA as a deterministic special case of a coupled latent variable model and, for the first time, integrate it into a variational autoencoding framework, yielding Var-JEPA. This approach introduces an explicit generative structure that unifies predictive and generative self-supervised learning, enabling meaningful representations without heuristic anti-collapse regularizers and supporting uncertainty quantification in the latent space. Leveraging ELBO optimization and a context–target prediction architecture, we instantiate Var-T-JEPA for tabular data, which significantly outperforms T-JEPA on downstream tasks and matches the performance of strong baseline methods using raw features.

generative modelingJoint-Embedding Predictive Architecturerepresentation learning

This work addresses the susceptibility of existing graph self-supervised learning methods to low-level input statistics and their limited capacity to model structural relationships among nodes. It introduces, for the first time, the Joint-Embedding Predictive Architecture (JEPA) paradigm to node-level graph representation learning through a structure-conditioned prediction mechanism: by masking k-hop subgraph structures, a context encoder predicts the target node’s representation in latent space, thereby circumventing reliance on input reconstruction or handcrafted augmentations. The approach integrates an EMA target encoder, cross-attention over spectral and centrality descriptors, and variance/covariance/Laplacian regularization, complemented by a progressive curriculum masking strategy to explicitly reinforce structural information learning. Evaluated on standard node classification benchmarks, the method achieves strong performance under both linear probing and fine-tuning, with ablation studies confirming the contribution of each component.

graph self-supervised learningjoint-embedding predictive architecturelatent prediction

CNN-JEPA: Self-Supervised Pretraining Convolutional Neural Networks Using Joint Embedding Predictive Architecture

Aug 14, 2024
AK
András Kalapos
🏛️ Budapest University of Technology and Economics

To address the incompatibility of the Joint Embedding Prediction Architecture (JEPA) paradigm with convolutional neural networks (CNNs), this paper proposes the first JEPA-based self-supervised learning framework specifically designed for CNNs. The method introduces three key innovations: (1) a sparse CNN encoder that supports masked inputs while preserving spatial sparsity; (2) a fully convolutional, depthwise-separable predictor that eliminates fully connected layers and auxiliary projection heads; and (3) an improved local masking strategy to enhance contextual modeling efficiency. Notably, the approach requires no explicit data augmentation. On ImageNet-100 with a ResNet-50 backbone, it achieves 73.3% top-1 linear evaluation accuracy, reduces training time by 17–35% compared to prior CNN-based JEPA methods, and matches or exceeds the linear and k-NN classification performance of leading contrastive and non-contrastive methods—including BYOL, SimCLR, and VICReg.

Adapts self-supervised learning to CNNs effectivelyImproves accuracy on ImageNet-100 with simpler architectureReduces training time significantly compared to other methods

Denoising with a Joint-Embedding Predictive Architecture

Oct 02, 2024
DC
Dengsheng Chen
🏛️ Meituan | Institute of Software, Chinese Academy of Sciences

This work addresses the unexplored limitations of Joint Embedding Predictive Architecture (JEPA) in generative modeling and introduces JEPA to generative tasks for the first time, proposing it as a unified framework for generalized token prediction. Methodologically, JEPA is reformulated as masked image modeling combined with continuous-space autoregressive denoising, enabling native compatibility with both diffusion and flow-matching losses and supporting multimodal continuous-data generation (e.g., video, audio). Key contributions include: (1) establishing the first JEPA-based generative paradigm; (2) deriving theoretical connections between JEPA and mainstream generative objectives—namely, diffusion and flow matching; and (3) achieving state-of-the-art performance on ImageNet conditional generation—demonstrating lower FID, faster convergence (reduced epochs), superior computational efficiency (notably scalable GFLOPs), and consistent optimality across baseline, large, and extra-large model scales.

Enhancing data generation flexibilityImproving ImageNet generation benchmarksIntegrating JEPA in generative modeling

Latest Papers

What's happening recently
View more

This study addresses the lack of visual intervention response modeling and counterfactual robustness analysis in I-JEPA by proposing CI-JEPA. The method introduces a counterfactual intervention-aware mechanism that quantifies semantic sensitivity and nuisance invariance by predicting representational shifts between original and modified images, thereby establishing a representation robustness evaluation framework grounded in selective sensitivity. Experimental evaluations employ joint embedding prediction, image masking interventions, and frozen-encoder linear probing techniques. Achieving 78.14% accuracy on the Flowers102 dataset, the results demonstrate that CI-JEPA significantly outperforms baseline models in relative semantic selectivity.

Counterfactual interventionJoint-embedding predictive architectureRepresentation robustness

This work addresses the limitations of existing Joint-Embedding Predictive Architectures (JEPAs), which typically employ a single encoder and consequently suffer from suboptimal representation learning efficiency and poor disentanglement. We propose SiamJEPA, a novel framework that introduces a masked Siamese student encoder paired with an exponential moving average (EMA) teacher network, enabling self-supervised learning through the prediction of latent embeddings of masked image regions. We demonstrate that the Siamese architecture is not merely a design choice but constitutes a crucial inductive bias that substantially enhances representational disentanglement and accelerates convergence during early training stages. Under linear probing on ImageNet with limited training budgets, SiamJEPA outperforms single-encoder JEPA variants and achieves higher accuracy than Masked Autoencoders (MAE), despite requiring significantly less training time.

inductive biasJoint Embedding Predictive Architecturepredictive learning

This work addresses the challenge that existing visual encoders struggle to recover semantic features of masked regions under severe occlusion, leading to a significant drop in classification performance. The authors propose leveraging the predictor trained within a Joint-Embedding Predictive Architecture (JEPA) as a transferable feature completion module. By applying a closed-form linear projection, this predictor can be seamlessly adapted to various non-JEPA encoders—such as CLIP—without any fine-tuning, enabling plug-and-play deployment. This study presents the first evidence that JEPA predictors can effectively transfer across distinct encoder families and supports fitting separate linear probes for varying mask ratios. Experiments demonstrate that integrating CLIP with the I-JEPA predictor boosts classification accuracy from 15.9% to 52.1% on heavily occluded ImageNet-9 and Stanford Dogs benchmarks, yielding a remarkable 36.2 percentage-point improvement.

feature recoveryheavy occlusionmasked representation learning

This work addresses the challenge of jointly modeling photometric invariance in images and temporal dynamics in videos within a unified framework. The authors propose UniJEPA, the first architecture that learns both image-level photometric prediction and video-level temporal state prediction end-to-end in a shared latent space, without relying on exponential moving averages (EMA), stop-gradient operations, or pretrained encoders. By combining next-embedding prediction loss with Gaussian regularization, UniJEPA achieves controllable abstraction: its photometric branch captures structural invariance, while its temporal branch learns dynamic equivariance. Experiments demonstrate that UniJEPA matches or exceeds the performance of specialized models across image, video, and control tasks, using only a single loss hyperparameter. Moreover, it enables zero-shot planning that is tens of times faster than generative world models while maintaining comparable accuracy.

Joint-Embedding Predictive Architecturelatent spaceself-supervised learning

Existing graph joint embedding prediction approaches rely on a single predefined partition, which limits their ability to capture multiscale structural information and consequently constrains the expressiveness and generalization of learned representations. This work proposes HP-JEPA, a novel framework that introduces hierarchical, multi-resolution graph partitioning into the JEPA architecture for the first time. HP-JEPA performs context-to-target latent prediction in parallel across multiple scales—from coarse to fine—and adaptively fuses these multiscale representations through task-aware weighting to jointly model local, regional, and global graph structures. Evaluated on eight benchmark datasets, HP-JEPA outperforms the fixed-resolution Graph-JEPA on six of them and demonstrates consistently superior performance across graphs of varying sizes, thereby validating the efficacy and advantages of the proposed multi-resolution strategy.

graph representation learninggraph self-supervised learninghierarchical partitioning

Hot Scholars

EA

Elie Aljalbout

Meta FAIR
RoboticsEmbodied AIMachine learningRobot learning
SH

Shashank Hegde

University of Southern California
machine learning