Scalable Patch-Level Self-Supervised Learning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of principled design in existing large-scale self-supervised learning methods, which typically rely on complex combinations of objectives. We propose the JEM algorithm, which formulates an information-theoretic objective grounded in a multi-view assumption. By integrating a student-teacher architecture with an information preservation loss and structure-preserving regularization, JEM achieves stable training through patch-level representation alignment. As the first latent-space patch-level method validated at the 7B parameter scale, it offers both theoretical rigor and engineering stability. Experimental results demonstrate that our approach outperforms DINOv2 on segmentation tasks. Furthermore, the 7B-parameter model surpasses the panoptic segmentation performance of DINOv3 while utilizing only one-twelfth of its training data.
📝 Abstract
Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on $12\times$ less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.
Problem

Research questions and friction points this paper is trying to address.

Self-supervised learning
Scalability
Patch-level representation
Information-theoretic objective
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Patch-Level Representation
Information-Theoretic Objective
Scalability
Student-Teacher Architecture