Geometry, Computation, and Optimality in Stochastic Optimization

📅 2019-09-23
📈 Citations: 10
Influential: 0
📄 PDF

career value

214K/year
🤖 AI Summary
This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.
📝 Abstract
We study computational and statistical consequences of problem geometry in stochastic and online optimization. By focusing on constraint set and gradient geometry, we characterize the problem families for which stochastic- and adaptive-gradient methods are (minimax) optimal and, conversely, when nonlinear updates -- such as those mirror descent employs -- are necessary for optimal convergence. When the constraint set is quadratically convex, diagonally pre-conditioned stochastic gradient methods are minimax optimal. We provide quantitative converses showing that the ``distance'' of the underlying constraints from quadratic convexity determines the sub-optimality of subgradient methods. These results apply, for example, to any $ell_p$-ball for $p<2$, and the computation/accuracy tradeoffs they demonstrate exhibit a striking analogy to those in Gaussian sequence models.
Problem

Research questions and friction points this paper is trying to address.

Characterize optimality of stochastic gradient methods via geometry
Determine when nonlinear updates are necessary for optimal convergence
Quantify sub-optimality of subgradient methods using constraint convexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diagonally pre-conditioned stochastic gradient methods
Characterizing problem families for optimal convergence
Quantifying sub-optimality via constraint geometry
C
Chen Cheng
Department of Statistics, Stanford University
J
John C. Duchi
Departments of Statistics and Electrical Engineering, Stanford University
Daniel Levy
Daniel Levy
National Heart, Lung, and Blood Institute
GeneticsGenomicsEpidemiologyCardiovascular Disease