Learning a Flow to Self-Supervised Representations

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of existing adversarial distribution matching methods, which hinders the efficient construction of geometry-referenced self-supervised representations. We propose FBDM, a non-adversarial flow-based distribution matching framework that learns geometric structures via spherical conditional velocity regression. By integrating an ETF-inspired reference with an explicit alignment loss, FBDM enables the number of reference components to exceed auxiliary dimensionality constraints and establishes theoretical bounds linking pretraining loss to downstream misclassification rates. Experiments spanning CIFAR to ImageNet benchmarks demonstrate that FBDM achieves performance comparable to adversarial approaches while accelerating training by 1.48–1.83× with minimal memory overhead. These results highlight its combined theoretical and empirical advantages for scalable representation learning.
📝 Abstract
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Problem

Research questions and friction points this paper is trying to address.

self-supervised learning
distribution matching
representation learning
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Based Distribution Matching
Self-Supervised Learning
Spherical Conditional Velocity Regression
Equiangular Tight Frame
Non-Adversarial Training