Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified theoretical foundation for stochastic gradient descent (SGD) convergence bounds and data shuffling strategies in high-dimensional machine learning. The proposed approach extends power-law spectral theory to stochastic optimization, demonstrating that the spectral exponent governs SGD dynamics. By integrating high-dimensional statistical learning theory with precise modeling of Gaussian data, tight convergence bounds featuring exact constants are derived. This work represents the first effort to unify abstract spectral theory with practical training sampling mechanisms, rigorously establishing the superiority of single shuffle over alternative sampling schemes. Ultimately, this research elucidates the intrinsic mechanism by which data geometry drives optimization speed, bridges the gap between theory and practice, and provides a comprehensive theoretical framework for stochastic training.
📝 Abstract
Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Optimization
Power-Law Spectra
Convergence Bounds
Data Shuffling
Stochastic Gradient Descent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Gradient Descent
Power-Law Spectra
Data Shuffling
Convergence Bounds
High-Dimensional Optimization