Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis
This study addresses the absence of a unified theoretical foundation for stochastic gradient descent (SGD) convergence bounds and data shuffling strategies in high-dimensional machine learning. The proposed approach extends power-law spectral theory to stochastic optimization, demonstrating that the spectral exponent governs SGD dynamics. By integrating high-dimensional statistical learning theory with precise modeling of Gaussian data, tight convergence bounds featuring exact constants are derived. This work represents the first effort to unify abstract spectral theory with practical training sampling mechanisms, rigorously establishing the superiority of single shuffle over alternative sampling schemes. Ultimately, this research elucidates the intrinsic mechanism by which data geometry drives optimization speed, bridges the gap between theory and practice, and provides a comprehensive theoretical framework for stochastic training.