๐ค AI Summary
This work addresses training instability and inefficiency in contrastive learning caused by excessively large gradient norms. We propose a spectral-domain analysis framework for gradient stability, establishing the first non-asymptotic spectral band constraint and proving an $O(1/ au^2)$ upper bound on the gradient normโrevealing the joint influence of batch spectral diversity, feature alignment, and temperature $ au$. Based on this theory, we design Greedy-64, a spectral-aware greedy batch selection algorithm that uses effective rank to quantify feature anisotropy for efficient batch construction. We further integrate batch whitening to suppress gradient variance. Experiments show that Greedy-64 accelerates training by 15% on ImageNet-100 while maintaining consistent accuracy gains on CIFAR-10; batch whitening reduces gradient variance by 1.37ร, empirically validating the theoretical bound.
๐ Abstract
We derive non-asymptotic spectral bands that bound the squared InfoNCE gradient norm via alignment, temperature, and batch spectrum, recovering the (1/ฯ^{2}) law and closely tracking batch-mean gradients on synthetic data and ImageNet. Using effective rank (R_{mathrm{eff}}) as an anisotropy proxy, we design spectrum-aware batch selection, including a fast greedy builder. On ImageNet-100, Greedy-64 cuts time-to-67.5% top-1 by 15% vs. random (24% vs. Pool--P3) at equal accuracy; CIFAR-10 shows similar gains. In-batch whitening promotes isotropy and reduces 50-step gradient variance by (1.37 imes), matching our theoretical upper bound.