🤖 AI Summary
This paper theoretically investigates the stability and generalization of minibatch SGD and local SGD, addressing a key gap in existing work—its overemphasis on optimization error while neglecting formal generalization guarantees. We propose a novel “training-error-driven stability analysis” framework, introducing for the first time an expectation-variance decomposition into stability modeling to explicitly characterize how training error influences generalization. Our theoretical analysis establishes that, under overparameterization, both algorithms achieve optimal generalization risk bounds; moreover, their generalization error decreases linearly with the number of parallel workers—a property we term *linear generalization speedup*. This result transcends conventional analyses focused solely on optimization convergence rates. To our knowledge, this is the first generalization theory for large-scale distributed learning that simultaneously provides rigorous stability guarantees and provable linear speedup in generalization performance.
📝 Abstract
The increasing scale of data propels the popularity of leveraging parallelism to speed up the optimization. Minibatch stochastic gradient descent (minibatch SGD) and local SGD are two popular methods for parallel optimization. The existing theoretical studies show a linear speedup of these methods with respect to the number of machines, which, however, is measured by optimization errors. As a comparison, the stability and generalization of these methods are much less studied. In this paper, we study the stability and generalization analysis of minibatch and local SGD to understand their learnability by introducing a novel expectation-variance decomposition. We incorporate training errors into the stability analysis, which shows how small training errors help generalization for overparameterized models. We show both minibatch and local SGD achieve a linear speedup to attain the optimal risk bounds.