Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the certifying capabilities and theoretical limitations of Gaussian regularizers in contrastive learning. Focusing on the shared Gaussianization approach, it proposes a characteristic function-based Gaussianity test that integrates χ_d-radius scaling, U-statistics, and rotation-invariant uniformity testing to systematically analyze detection mechanisms for view misalignment and non-uniformity. The main contributions include proving that the InfoNCE excess loss is tightly bounded by the square-root rate of the SG loss, thereby revealing how alignment weights influence local minima. Furthermore, this work establishes dimension-free sharp constant bounds, validates the effectiveness of specific Gaussian kernels, and quantifies the impact of batch size on critical weights. Collectively, these results provide a rigorous theoretical foundation for the LeJEPA framework.
📝 Abstract
What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χ_d$ radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by $4\cdot 3^{3/4}β$ times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{βu^\top v}$; SG's own test does, Gaussian kernels $e^{-γ\|u-v\|^2}$ qualify exactly when $γ\ge β/2$, and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel's gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG$_{0.2}$ loss than SG$_{0.2}$ training from scratch, which points to an optimization gap.
Problem

Research questions and friction points this paper is trying to address.

contrastive learning
distribution matching
shared Gaussianization
InfoNCE
regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shared Gaussianization
Contrastive Learning
InfoNCE
Distribution-matching Regularizer
LeJEPA
🔎 Similar Papers
No similar papers found.