joint-embedding predictive learning

Design and implement self-supervised architectures and training procedures that learn a shared embedding space by predicting latent target features of masked or alternative views (within or across modalities), using an encoder, target encoder, and predictor to enforce a single joint-embedding predictive objective. Build models and loss schemes that preserve within-modal neighborhood relations, support a single shared encoder or multi-modal inputs, and scale to large datasets.

joint-embeddingpredictivelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning

May 18, 2025
HV
Hugues Van Assel
🏛️ Genentech | Meta AI | Brown University

The lack of principled criteria for choosing between contrastive joint-embedding and generative reconstruction paradigms in self-supervised learning (SSL) hinders theoretical understanding and practical design. Method: We conduct a rigorous theoretical analysis under linear model assumptions, deriving closed-form solutions for both paradigms and explicitly modeling the view-generation process to characterize the impact of data augmentations and nuisance features on representation learning. Results: Our analysis reveals that joint-embedding achieves asymptotically optimal performance under strong nuisance features with weaker alignment requirements, whereas reconstruction is inherently sensitive to such features. We further derive the minimal necessary condition linking augmentation strength and feature alignment. This work provides the first quantitative, interpretable explanation for the empirical superiority of joint-embedding over reconstruction on complex real-world data, establishing the first theoretically grounded, explainable guidance for SSL paradigm selection.

Analyze impact of view generation on representationsCompare reconstruction and joint embedding in SSLDetermine optimal SSL paradigm for irrelevant features

This work addresses the limitations of existing audio-visual self-supervised learning methods, which rely on modality-specific encoders and complex objective functions that hinder effective cross-modal synergy. The authors propose the first Joint Embedding Predictive Architecture (JEPA) for audio-visual representation learning, featuring a modality-agnostic unified encoder and a single predictive objective that jointly models intra- and inter-modal relationships to enable complementary information exchange across modalities. With a frozen ViT-g backbone, the method surpasses the previous best frozen baseline by 6.8 mAP on AudioSet-20K and outperforms fully fine-tuned models on ESC-50 and FSD50K. Notably, it achieves competitive performance on video tasks using only one-tenth of the video data, demonstrating substantially improved representation efficiency and generalization capability.

audio-visual learningcross-modal representationjoint embedding

Graph-level Representation Learning with Joint-Embedding Predictive Architectures

Sep 27, 2023
GS
Geri Skenderi
🏛️ Bocconi University | Michigan State University | University of Verona

To address the limitations of contrastive learning—such as susceptibility to overfitting and difficulty in capturing semantic hierarchies—in graph-level representation learning, this paper proposes Graph-JEPA, the first adaptation of the Joint Embedding Predictive Architecture (JEPA) to graph-structured data. Graph-JEPA enables contrastive-free and reconstruction-free self-supervision by masking subgraphs and predicting their latent representations. Crucially, it introduces hyperbolic coordinate regression as a novel objective to explicitly model the implicit hierarchical structure among graph concepts. By eliminating negative sampling and pixel-level reconstruction, Graph-JEPA significantly mitigates overfitting. Extensive experiments demonstrate that Graph-JEPA consistently outperforms state-of-the-art self-supervised methods on graph classification, continuous-value regression, and non-isomorphic graph discrimination tasks. The learned graph-level representations exhibit superior semantic richness and generalization capability.

Graph FeaturesMachine LearningSelf-supervised Learning

This work proposes Predictive Representation Learning (PRL), a novel paradigm in self-supervised learning that moves beyond conventional approaches limited to representation alignment and input reconstruction, which often fail to model unobserved regions of the data distribution. The study formally defines the PRL framework for the first time and identifies the Joint-Embedding Predictive Architecture (JEPA) as its canonical instantiation, thereby establishing a unified taxonomy encompassing alignment, reconstruction, and prediction. Comparative experiments with BYOL, MAE, and I-JEPA demonstrate that PRL methods achieve high accuracy (BYOL: 0.98; I-JEPA: 0.95) while significantly enhancing robustness (0.75 and 0.78, respectively). In contrast, purely reconstructive approaches like MAE attain perfect similarity (1.00) but exhibit markedly lower robustness (0.55), underscoring the critical role of predictive mechanisms in improving representational generalization.

data distribution predictionlatent predictionpredictive representation learning

This work addresses the limitations of traditional generative modeling, which often focuses on pixel-level reconstruction and struggles to capture high-level semantics. To overcome this, the authors propose an Energy-Based Joint Embedding Predictive Architecture (EB-JEPA) that performs self-supervised prediction in representation space rather than pixel space, effectively enabling the construction of world models for images, videos, and action-conditioned environments. The study introduces the first open-source, lightweight, and modular EB-JEPA library, systematically demonstrating the critical role of regularization in preventing representational collapse. The framework supports multi-step temporal prediction and action-conditioned modeling, achieving strong empirical results: 91% probe accuracy on CIFAR-10, high-quality multi-step video prediction on Moving MNIST, and a 97% planning success rate on the Two Rooms navigation task—all trained within hours on a single GPU.

Energy-Based ModelsJoint-Embedding Predictive ArchitecturesRepresentation Learning

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified theoretical foundation in unsupervised visual representation learning, where existing methods struggle to simultaneously achieve semantic invariance, spatial structure modeling, and non-degenerate solutions. The authors propose three essential principles—observation, prediction, and regularization—and formulate them within a unified energy-based decomposition framework, offering the first formalization of core self-supervised learning criteria. Through rigorous analysis of gradient complementarity, convergence guarantees for momentum encoders, and a negative-sample-free alignment theory, the study exposes fundamental limitations of contrastive learning and momentum mechanisms, demonstrating that all three principles are indispensable. Controlled experiments, including block retrieval evaluations, confirm that optimal performance is attained only when these principles operate in concert.

non-degeneracyself-supervised learningsemantic invariance

This work addresses the challenge of jointly modeling photometric invariance in images and temporal dynamics in videos within a unified framework. The authors propose UniJEPA, the first architecture that learns both image-level photometric prediction and video-level temporal state prediction end-to-end in a shared latent space, without relying on exponential moving averages (EMA), stop-gradient operations, or pretrained encoders. By combining next-embedding prediction loss with Gaussian regularization, UniJEPA achieves controllable abstraction: its photometric branch captures structural invariance, while its temporal branch learns dynamic equivariance. Experiments demonstrate that UniJEPA matches or exceeds the performance of specialized models across image, video, and control tasks, using only a single loss hyperparameter. Moreover, it enables zero-shot planning that is tens of times faster than generative world models while maintaining comparable accuracy.

Joint-Embedding Predictive Architecturelatent spaceself-supervised learning

This work identifies and formally names a previously uncharacterized phenomenon in vision-language models termed “cross-modal feature heterogeneity,” wherein semantically equivalent concepts activate sparse features along inconsistent directions across modalities, leading to modality fragmentation that undermines interpretability and controllability. The study demonstrates that mere alignment of activations is insufficient to resolve this feature mismatch. To address this, the authors propose a two-stage strategy: first preserving each modality’s intrinsic feature geometry using modality-specific sparse autoencoders, followed by post-hoc alignment of corresponding cross-modal features. This approach significantly improves reconstruction fidelity and achieves superior performance in cross-modal retrieval and concept-guided generation tasks.

cross-modal feature heterogeneityfeature alignmentmodality split

This study systematically investigates the interplay between self-supervised and supervised learning objectives by comparing the pretrain-then-finetune (PFT) and joint training (JT) paradigms across varying label budgets, task types, and data domains. For the first time, it conducts a comprehensive evaluation of representation quality, robustness, and cross-domain generalization using eight representative self-supervised methods across diverse visual domains—including natural images, medical imaging, emergency response, and remote sensing. The findings reveal that JT is more efficient and robust under low-label regimes, whereas PFT demonstrates greater reliability in specialized, complex domains. These results provide empirical guidance for selecting training strategies in practical applications and establish a new benchmark for hybrid self-supervised semi-supervised learning.

Joint TrainingPretrain-FinetuningSelf-Supervised Learning

Existing multimodal image clustering methods often compromise modality-specific structures by directly aligning cross-modal representations, leading to unreliable alignment. To address this issue, this work proposes DeepMORSE, which introduces a novel modality-shared self-expressive mechanism that preserves the intrinsic subspace structure of each modality while enabling effective cross-modal alignment. Theoretically, this mechanism suppresses inter-class noise and promotes subspace-preserving solutions, with mini-batch optimization implicitly inducing a regularizing effect. Leveraging deep representations from vision-language models and self-expressive learning, DeepMORSE achieves over 3% absolute improvement in clustering performance across six benchmarks—including UCF-101, DTD-47, and ImageNet-Dogs—and attains state-of-the-art results in image retrieval and zero-shot classification without task-specific losses or post-processing.

cross-modal alignmentimage clusteringmodality-specific structure

Hot Scholars

FS

Fahad Shahbaz Khan

MBZUAI, Linköping University Sweden
Computer VisionObject RecognitionGenerative AIAI for Science
TG

Theo Gevers

Computer Vision Research Group, University of Amsterdam (UvA); 3DUniversum
computer visionimage understandingdeep learningobject recognition
MP

Marcin Przewięźlikowski

Jagiellonian University, GMUM
Machine LearningDeep LearningImage RecognitionFew-Shot Learning
BZ

Bin Zhu

Assistant Professor, Singapore Management University
MultimediaComputer Vision
CB

Chenxi Bao

MBZUAI
Music GenerationInteractive Music DesignComputer Music