Two Global Crops Suffice: Locating Semantic Emergence in DINO-Style Self-Supervised Learning

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过几何不同全局视图的一致性解决了DINO风格自监督学习中的语义表示质量问题。
📝 Abstract
Self-supervised vision transformers trained with DINO-style objectives exhibit striking emergent semantic representation quality across visual tasks, yet the mechanisms underlying this behavior remain unclear. We present a systematic empirical dissection of the DINO family and show that semantic representations arise primarily from enforcing consistency between geometrically distinct global views of the same image instance. This instance-specific global alignment acts as the semantic anchor of DINO-style learning. Across controlled retraining experiments evaluated on semantic correspondence and a diverse suite of 2D and 3D downstream tasks, we find that patch-level masking objectives enhance semantics only when trained jointly with this global alignment, indicating that the iBOT objective refines and densifies existing semantic structure rather than creating it independently. In contrast, local-to-global view alignment does not substantially improve semantic qualities at fixed compute beyond a purely global alignment. Beyond training design, we revisit how semantic representation quality should be evaluated: while classification accuracy is the standard validation score, semantic correspondence provides a complementary axis that more reliably predicts downstream task performance. Together, these findings provide a functional decomposition of DINO-style learning and represent an important step toward understanding how semantic representations emerge in self-supervised vision models.
Problem

Research questions and friction points this paper is trying to address.

DINO-style learning
semantic representation quality
self-supervised vision transformers
emergent behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

global alignment
semantic emergence
self-supervised learning
DINO-style objectives
semantic correspondence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Basavaraj Sunagad
CISPA Helmholtz Center for Information Security
A
Artur Jesslen
University of Freiburg
Adam Kortylewski
Adam Kortylewski
Research Group Leader, University of Freiburg and Max Planck Institute for Informatics
Visual ComputingMachine LearningGenerative AI