🤖 AI Summary
This study addresses the issues of representational dimensional collapse and limited transferability arising from the mismatch between projection heads and backbone networks in joint-embedding self-supervised learning. To this end, we propose SACReg, a spectral anti-collapse regularization method. Leveraging theoretical derivations based on linear network theory and spectral analysis techniques, we reveal the mechanism underlying low-rank collapse and introduce a novel regularization strategy grounded in λ-balance theory to effectively prevent representational degradation in the backbone. Extensive evaluations demonstrate that SACReg significantly outperforms baselines such as LeJEPA and V-JEPA 2 across ImageNet classification, multi-dataset linear probing, and video understanding benchmarks, comprehensively enhancing both the representation quality and downstream generalization performance of JEPA models.
📝 Abstract
Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $\lambda$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $\lambda$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $\lambda$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations'ranks. We apply SACReg to JEPA and propose $\lambda$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $\lambda$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.