🤖 AI Summary
To address the incompatibility of the Joint Embedding Prediction Architecture (JEPA) paradigm with convolutional neural networks (CNNs), this paper proposes the first JEPA-based self-supervised learning framework specifically designed for CNNs. The method introduces three key innovations: (1) a sparse CNN encoder that supports masked inputs while preserving spatial sparsity; (2) a fully convolutional, depthwise-separable predictor that eliminates fully connected layers and auxiliary projection heads; and (3) an improved local masking strategy to enhance contextual modeling efficiency. Notably, the approach requires no explicit data augmentation. On ImageNet-100 with a ResNet-50 backbone, it achieves 73.3% top-1 linear evaluation accuracy, reduces training time by 17–35% compared to prior CNN-based JEPA methods, and matches or exceeds the linear and k-NN classification performance of leading contrastive and non-contrastive methods—including BYOL, SimCLR, and VICReg.
📝 Abstract
Self-supervised learning (SSL) has become an im-portant approach in pretraining large neural networks, enabling unprecedented scaling of model and dataset sizes. While recent advances like I-JEPA have shown promising results for Vision Transformers, adapting such methods to Convolutional Neural Networks (CNNs) presents unique challenges. In this paper, we introduce CNN -JEPA, a novel SSL method that successfully applies the joint embedding predictive architecture approach to CNN s. Our method incorporates a sparse CNN encoder to handle masked inputs, a fully convolutional predictor using depthwise separable convolutions, and an improved masking strategy. We demonstrate that CNN-JEPA outperforms I-JEPA with ViT architectures on ImageNet-l00, achieving a 73.3% linear top-1 accuracy using a standard ResN et-50 encoder. Compared to other CNN-based SSL methods, CNN-JEPA requires 17– 35 % less training time for the same number of epochs and approaches the linear and k-NN top-l accuracies of BYOL, SimCLR, and VICReg. Our approach offers a simpler, more efficient alternative to existing SSL methods for CNNs, requiring minimal augmentations and no separate projector network.