The Adaptive Complexity of Finding a Stationary Point

📅 2025-05-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies the *adaptive parallel complexity* of finding an ε-stationary point in nonconvex optimization—i.e., the minimal number of sequential rounds required to achieve ε-stationarity, assuming polynomially many oracle queries per round. We establish *tight upper and lower bounds* on this complexity in both high-dimensional and constant-dimensional settings. In high dimensions, we prove that gradient descent, cubic-regularized Newton’s method, and p-th-order adaptive regularization methods all achieve the optimal round complexity Ω(ε^{−(p+1)/p}). In constant dimensions, we propose a novel algorithm combining grid search with gradient-flow trapping, attaining O(1) rounds. We further derive a per-round query lower bound of Ω̃(ε^{−(d−1)/2}), confirming its adaptive optimality. Key technical tools include chain-based hardness potential analysis, Lipschitz p-th-order derivative constructions, gradient-flow trapping phenomena, and information-theoretic lower-bound derivation.

Technology Category

Search and Optimization: Non-convex OptimizationMachine Learning: OptimizationReasoning under Uncertainty: Stochastic Optimization

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSecurity and Privacy: Large-scale security measurements
📝 Abstract
In large-scale applications, such as machine learning, it is desirable to design non-convex optimization algorithms with a high degree of parallelization. In this work, we study the adaptive complexity of finding a stationary point, which is the minimal number of sequential rounds required to achieve stationarity given polynomially many queries executed in parallel at each round. For the high-dimensional case, i.e., $d = widetilde{Omega}(varepsilon^{-(2 + 2p)/p})$, we show that for any (potentially randomized) algorithm, there exists a function with Lipschitz $p$-th order derivatives such that the algorithm requires at least $varepsilon^{-(p+1)/p}$ iterations to find an $varepsilon$-stationary point. Our lower bounds are tight and show that even with $mathrm{poly}(d)$ queries per iteration, no algorithm has better convergence rate than those achievable with one-query-per-round algorithms. In other words, gradient descent, the cubic-regularized Newton's method, and the $p$-th order adaptive regularization method are adaptively optimal. Our proof relies upon novel analysis with the characterization of the output for the hardness potentials based on a chain-like structure with random partition. For the constant-dimensional case, i.e., $d = Theta(1)$, we propose an algorithm that bridges grid search and gradient flow trapping, finding an approximate stationary point in constant iterations. Its asymptotic tightness is verified by a new lower bound on the required queries per iteration. We show there exists a smooth function such that any algorithm running with $Theta(log (1/varepsilon))$ rounds requires at least $widetilde{Omega}((1/varepsilon)^{(d-1)/2})$ queries per round. This lower bound is tight up to a logarithmic factor, and implies that the gradient flow trapping is adaptively optimal.
Problem

Research questions and friction points this paper is trying to address.

Study adaptive complexity for finding stationary points in non-convex optimization
Prove tight lower bounds for high-dimensional optimization algorithms
Develop optimal algorithm for constant-dimensional stationary point search
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-convex optimization with high parallelization
Lower bounds for adaptive complexity in high dimensions
Algorithm bridging grid search and gradient flow
🔎 Similar Papers
💼 Related Jobs
No related jobs found.