Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes

๐Ÿ“… 2025-10-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses training instability and inefficiency in contrastive learning caused by excessively large gradient norms. We propose a spectral-domain analysis framework for gradient stability, establishing the first non-asymptotic spectral band constraint and proving an $O(1/ au^2)$ upper bound on the gradient normโ€”revealing the joint influence of batch spectral diversity, feature alignment, and temperature $ au$. Based on this theory, we design Greedy-64, a spectral-aware greedy batch selection algorithm that uses effective rank to quantify feature anisotropy for efficient batch construction. We further integrate batch whitening to suppress gradient variance. Experiments show that Greedy-64 accelerates training by 15% on ImageNet-100 while maintaining consistent accuracy gains on CIFAR-10; batch whitening reduces gradient variance by 1.37ร—, empirically validating the theoretical bound.

Technology Category

Machine Learning: Learning with ManifoldsComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web dataSecurity and Privacy: Large-scale security measurements
๐Ÿ“ Abstract
We derive non-asymptotic spectral bands that bound the squared InfoNCE gradient norm via alignment, temperature, and batch spectrum, recovering the (1/ฯ„^{2}) law and closely tracking batch-mean gradients on synthetic data and ImageNet. Using effective rank (R_{mathrm{eff}}) as an anisotropy proxy, we design spectrum-aware batch selection, including a fast greedy builder. On ImageNet-100, Greedy-64 cuts time-to-67.5% top-1 by 15% vs. random (24% vs. Pool--P3) at equal accuracy; CIFAR-10 shows similar gains. In-batch whitening promotes isotropy and reduces 50-step gradient variance by (1.37 imes), matching our theoretical upper bound.
Problem

Research questions and friction points this paper is trying to address.

Derives spectral bounds for InfoNCE gradient norms in contrastive learning
Designs spectrum-aware batch selection to accelerate training convergence
Uses in-batch whitening to reduce gradient variance and promote isotropy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Derives spectral bounds on InfoNCE gradient norms
Designs spectrum-aware batch selection using effective rank
Uses in-batch whitening to reduce gradient variance
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
P
Peter Ochieng
Department of Computer Science, University of Cambridge