Which Tasks Survive Self-Supervised Learning?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms underlying downstream task information preservation in self-supervised learning (SSL) and investigates the relationship between semantic recoverability and representation geometry. We propose a theoretical framework for semantic recoverability, proving its alignment with directional geometry and deriving a closed-form characterization based on spectral properties. The framework is validated through linear probing, centroid axis analysis, and spectral decomposition of two-view operators on both synthetic and real-world datasets. Our contributions establish theoretical connections among recoverability, directional geometry, and few-shot transfer, while identifying the spectral conditions governing task information retention. Experiments confirm that recoverability plays a decisive role in the normalized variance of class distances and few-shot classification performance, offering new perspectives on understanding information preservation in SSL.
📝 Abstract
Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation. We study this question through \emph{semantic recoverability}, defined as the amount of a task's posterior score captured by the represented function space. We show that, for centered and whitened representations, recoverability exactly determines directional class-distance-normalized variance (CDNV), controls few-shot nearest-centroid classification, and governs the strength of task-relevant semantic directions. The population linear probe and centroid axis coincide, and multiple well-recovered tasks approach a factorial centroid geometry. We then analyze a canonical two-view SSL objective and show that its population optimum spans the leading cross-view-stable modes of the associated two-view operator. This yields a closed-form spectral characterization of semantic recoverability: a downstream task is preserved to the extent that its posterior lies in the selected spectral subspace. We validate these predictions on synthetic and real datasets across several SSL methods, testing the predicted relationships among recoverability, directional geometry, spectral structure, and few-shot transfer. Together, these results give a task-level account of what information survives same-instance SSL and how the retained information appears in downstream geometry and transfer.
Problem

Research questions and friction points this paper is trying to address.

Self-Supervised Learning
Semantic Recoverability
Downstream Tasks
Representation Learning
Spectral Characterization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Semantic Recoverability
Spectral Characterization
Few-Shot Transfer
Representation Geometry
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Achleshwar Luthra
Achleshwar Luthra
Texas A&M University
Deep LearningComputer Vision
L
Lucas Bryant
Department of Computer Science and Engineering, Texas A&M University
T
Tracy Zhu
Department of Computer Science and Engineering, Texas A&M University
Tomer Galanti
Tomer Galanti
AI Researcher
Deep LearningMachine Learning