Convex Optimization Is Free When Accuracy Is Expensive

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the efficiency bottleneck in convex optimization arising from the prohibitive cost of gradient approximation within higher-order momentum control (HTMC) scenarios. To overcome this limitation, we propose a randomized multi-level oracle mechanism that transforms deterministic approximations into unbiased estimators, thereby enabling independent pricing of accuracy and variance. We theoretically demonstrate that the optimization cost can be reduced to the level of a single gradient evaluation, revealing that this cost is fundamentally a functional of the gradient flow rather than an artifact of discretization. Building upon randomized multi-level sampling and inexact gradient descent algorithms, our approach reduces the computational complexity to O(Ρ⁻ᡞ) for convex problems and O(Ρ⁻ᡞ/²) for strongly convex problems, surpassing the limitations of conventional methods by an entire order of magnitude.
πŸ“ Abstract
This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $\delta^{-\gamma}$ in the accuracy $\delta$. When $\gamma>2$, falling into the Harder-Than-Monte-Carlo (HTMC) regime, the price of accuracy outruns the variance reduction that Monte Carlo would buy and we show that minimizing a loss function costs no more, up to a factor depending only on $\gamma$, than a single evaluation of its gradient at the accuracy the problem demands. A randomized multilevel oracle replaces the deterministic approximation of accuracy $\delta$ by an unbiased estimator of it, whose variance $\sigma^2$ becomes a second, independently priced dial: the cost of one call drops from $\delta^{-\gamma}$ to $\delta^{2-\gamma}\sigma^{-2}$. Plain inexact gradient descent driven by that oracle reaches loss $\varepsilon$ at expected compute $\Theta(\varepsilon^{-\gamma})$ in the convex case, against $\Theta(\varepsilon^{-(\gamma+1)})$ for the same method run at a fixed accuracy: randomization buys a full power of $\varepsilon$. Under $\mu$-strong convexity the exponent halves, to $\varepsilon^{-\gamma/2}$, because the iterates settle at a noise floor and the bias budget relaxes accordingly. Both bounds are independent of the step size, and hence of the smoothness constant, and we show that the cost is a functional of the underlying gradient flow rather than of any discretization of it.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Convex Optimization
Multilevel Oracle
Inexact Gradient Descent
Computational Complexity
Variance Reduction