Why Self-Supervised Encoders Want to Be Normal

📅 2026-04-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates theoretical guarantees for representation learning in self-supervised and semi-supervised settings, aiming to balance information compression with predictive power. Framed through the information bottleneck principle, the problem is cast as a rate–distortion optimization, where optimal representations are obtained via soft clustering on a predictive manifold. The authors propose Sketched Isotropic Gaussian Regularization (SIGReg), which constructs an exact transformation chain from the probability simplex to an isotropic Gaussian distribution, yielding a tractable, non-variational encoder loss. Theoretical analysis combines conditional entropy bottleneck decomposition with minibatch-based marginal estimation. Empirical validation on synthetic data and FashionMNIST demonstrates the effectiveness of the rate–distortion trade-off, with the non-parametric implementation achieving performance comparable to standard variational methods.
📝 Abstract
We develop a geometric and information-theoretic framework for encoder-decoder learning built on the Information Bottleneck (IB) principle. Recasting IB as a rate-distortion problem with Kullback-Leibler (KL) divergence as distortion, we show that the optimal representation at any distortion level is a soft clustering of the \emph{predictive manifold} $\mathcal{M}=\{p(Y|x):x\in\mathcal{X}\}$ inside the probability simplex, admitting a linear decoder in the canonical parameterization. We derive a chain of exact transformations, from flat Dirichlet to exponential to isotropic Gaussian, connecting the maximum entropy prior on the simplex to Euclidean space, with quantified entropy overhead at each step, and show that Sketched Isotropic Gaussian Regularization (SIGReg) implements a Gaussian relaxation of this principle whose overhead affects rate accounting but not achievable prediction. This relaxation provides a principled distributional regularizer for learning with limited or no supervision. Using the Conditional Entropy Bottleneck (CEB) decomposition, we derive concrete encoder losses for supervised and semi-supervised settings, estimated via minibatch marginals without variational bounds. In the self-supervised setting, the CEB conditional rate is replaced by a view-prediction proxy. SIGReg serves as the distributional regularizer for both the semi-supervised and self-supervised settings. Experiments on toy problems and FashionMNIST confirm the predicted rate-distortion trade-offs and show that the non-parametric estimator is competitive with the standard variational approach.
Problem

Research questions and friction points this paper is trying to address.

Self-Supervised Learning
Information Bottleneck
Representation Learning
Distributional Regularization
Rate-Distortion Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Information Bottleneck
Rate-Distortion Theory
SIGReg
Predictive Manifold
Self-Supervised Learning
🔎 Similar Papers
2024-05-29arXiv.orgCitations: 0