CNN-JEPA: Self-Supervised Pretraining Convolutional Neural Networks Using Joint Embedding Predictive Architecture

📅 2024-08-14
🏛️ International Conference on Machine Learning and Applications
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the incompatibility of the Joint Embedding Prediction Architecture (JEPA) paradigm with convolutional neural networks (CNNs), this paper proposes the first JEPA-based self-supervised learning framework specifically designed for CNNs. The method introduces three key innovations: (1) a sparse CNN encoder that supports masked inputs while preserving spatial sparsity; (2) a fully convolutional, depthwise-separable predictor that eliminates fully connected layers and auxiliary projection heads; and (3) an improved local masking strategy to enhance contextual modeling efficiency. Notably, the approach requires no explicit data augmentation. On ImageNet-100 with a ResNet-50 backbone, it achieves 73.3% top-1 linear evaluation accuracy, reduces training time by 17–35% compared to prior CNN-based JEPA methods, and matches or exceeds the linear and k-NN classification performance of leading contrastive and non-contrastive methods—including BYOL, SimCLR, and VICReg.

Technology Category

Machine Learning: Unsupervised & Self-Supervised LearningComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Self-supervised learning (SSL) has become an im-portant approach in pretraining large neural networks, enabling unprecedented scaling of model and dataset sizes. While recent advances like I-JEPA have shown promising results for Vision Transformers, adapting such methods to Convolutional Neural Networks (CNNs) presents unique challenges. In this paper, we introduce CNN -JEPA, a novel SSL method that successfully applies the joint embedding predictive architecture approach to CNN s. Our method incorporates a sparse CNN encoder to handle masked inputs, a fully convolutional predictor using depthwise separable convolutions, and an improved masking strategy. We demonstrate that CNN-JEPA outperforms I-JEPA with ViT architectures on ImageNet-l00, achieving a 73.3% linear top-1 accuracy using a standard ResN et-50 encoder. Compared to other CNN-based SSL methods, CNN-JEPA requires 17– 35 % less training time for the same number of epochs and approaches the linear and k-NN top-l accuracies of BYOL, SimCLR, and VICReg. Our approach offers a simpler, more efficient alternative to existing SSL methods for CNNs, requiring minimal augmentations and no separate projector network.
Problem

Research questions and friction points this paper is trying to address.

Adapts self-supervised learning to CNNs effectively
Reduces training time significantly compared to other methods
Improves accuracy on ImageNet-100 with simpler architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse CNN encoder for masked inputs
Fully convolutional predictor with depthwise convolutions
Improved masking strategy for efficient training
🔎 Similar Papers
2024-03-07arXiv.orgCitations: 2
Budapest University of Technology and Economics
A
András Kalapos
Department of Telecommunications and Artificial Intelligence, Faculty of Electrical Engineering and Informatics, Budapest University of Technology and Economics
B
B'alint Gyires-T'oth
Department of Telecommunications and Artificial Intelligence, Faculty of Electrical Engineering and Informatics, Budapest University of Technology and Economics