Score
Establish and apply the Kurdyka-Łojasiewicz (KL) property and KL inequality for objective functions to analyze iterative optimization methods; this includes proving a function satisfies the KL property and using descent/relative-error conditions to show finite-length iterates, convergence to critical points, and explicit convergence rates (including for nonsmooth or adaptive algorithms).
This paper studies the oracle complexity of finding a (δ,ε)-stable point—a point whose δ-neighborhood contains a subgradient of norm at most ε—in nonsmooth optimization. Under Lipschitz continuity, it establishes for the first time that deterministic first-order algorithms necessarily incur dimension-dependent complexity, precluding dimension-free bounds; in contrast, randomized algorithms achieve the tight upper bound Õ(1/(δε³)), matched by a universal randomized lower bound. It further reveals that convexity dramatically accelerates convergence: for convex functions, a deterministic O(1/ε²) upper bound is attained, and it is proven that zero subgradients cannot be identified exactly in finite time; for smooth functions, derandomization is achievable with only logarithmic overhead. The core contribution lies in precisely characterizing the fundamental roles of determinism vs. randomness, convexity, and smoothness in governing oracle complexity, and providing tight upper and lower bounds for each setting.
This work addresses the open problem posed by Syrgkanis et al. concerning last-iterate convergence of Optimistic Multiplicative Weights Update (OMWU) in constrained convex-concave minimax optimization—particularly relevant to zero-sum games and GANs. We establish, for the first time, global convergence of OMWU to exact saddle points under general convex constraints, without requiring unconstrained domains or strong regularization assumptions. Our analysis introduces a novel framework combining monotonic KL-divergence descent with local contraction mapping properties, integrating fixed-point theory and contraction mapping techniques. This approach overcomes key limitations of prior analyses reliant on either unbounded domains or stringent regularization. The result provides the first rigorous theoretical guarantee for OMWU’s last-iterate convergence in constrained saddle-point optimization and furnishes a principled foundation for termination criteria in practical iterative implementations.
This paper addresses strongly convex yet nonsmooth and non-Lipschitz optimization problems. We develop a unified primal–dual theoretical framework that, for the first time, reveals the equivalent dual-averaging representations of the subgradient method, proximal subgradient method, and switching subgradient method. Through a novel *dual-gap convergence analysis*, we establish the first $O(1/T)$ convergence guarantee applicable to this problem class. We derive an optimal stopping criterion and optimality certificate that require no additional computation, and rigorously characterize a controllable convergence boundary—even under early-stage exponential divergence. Our theory accommodates a broad range of step-size choices and accommodates ill-conditioned non-Lipschitz structures. While preserving algorithmic simplicity, our framework substantially extends both the applicability and theoretical depth of subgradient-type methods.
This work investigates the theoretical hardness and algorithmic efficiency of finding stationary points of the hyperobjective in nonconvex–convex and nonconvex–nonconvex bilevel optimization, under the Polyak–Łojasiewicz (PL) condition—rather than strong convexity—on the lower-level objective. We first establish an impossibility result: for zero-respecting algorithms, computing a hyperstationary point is fundamentally intractable in the nonconvex–convex setting. Under the PL condition, we break the reliance on strong convexity and derive tighter hypergradient convergence complexity bounds. We propose a novel analytical framework unifying implicit function differentiation, hypergradient estimation, and first-order optimization. This yields complexity guarantees of $ ilde{mathcal{O}}(varepsilon^{-2})$, $ ilde{mathcal{O}}(varepsilon^{-4})$, and $ ilde{mathcal{O}}(varepsilon^{-6})$ for deterministic, partially stochastic, and fully stochastic settings, respectively—substantially improving upon existing nonconvex bilevel optimization methods.
For nonsmooth weakly convex optimization, existing methods struggle to escape strict saddle points—hindering convergence to local minima. Method: This paper proposes a family of perturbed proximal algorithms—including perturbed proximal point, proximal gradient, and proximal linear variants—to address this challenge. Contributions/Results: We establish, for the first time, a verifiable characterization of ε-approximate local minima for nonsmooth weakly convex functions and provide the first theoretical guarantee for escaping strict saddle points. By integrating perturbed optimization, nonsmooth analysis, and saddle-point escape theory, all three algorithms achieve an iteration complexity of O(ε⁻² log d) under standard assumptions to compute an ε-approximate local minimum. This yields the first polynomial-time convergence guarantee for escaping saddle points in nonsmooth optimization.
This work addresses nonconvex, nonsmooth stochastic optimization problems where the objective comprises an expected loss and a lower semicontinuous regularizer, without assuming convexity or uniformly bounded gradient variance. The authors propose a proximal stochastic subgradient method incorporating an Armijo-type line search, which leverages progressively refined sample average approximations. Under the mild condition that the sample size is nondecreasing and diverges to infinity, they establish—within the nonsmooth Kurdyka–Łojasiewicz (KL) framework—the almost sure convergence of function values, global trajectory convergence, and the property that all cluster points are stationary. Notably, this is the first such result in the nonsmooth KL setting. Furthermore, for functions satisfying a KL inequality with exponential desingularizing functions, they derive a polynomial convergence rate up to a logarithmic factor, thereby relaxing conventional requirements on regularizer convexity and variance boundedness.
This paper investigates the convergence of gradient-free optimization algorithms under noisy, nonsmooth, or both settings—typical in black-box optimization where gradients are unavailable. We propose a unified analytical framework integrating model-based strategies and smoothing techniques, replacing gradient estimation for nonsmooth or stochastic objectives with smooth approximations via a generalized gradient descent recursion. Under minimal regularity assumptions—requiring only local bounded variation—we rigorously establish convergence guarantees for both deterministic and stochastic settings. Our analysis uncovers a fundamental trade-off between the smoothing parameter and step size, characterizing their joint impact on convergence rate and stability. Extensive experiments on diverse machine learning classification tasks demonstrate the method’s effectiveness and robustness against noise and nonsmoothness.
This study addresses the minimization of convex functions that are relatively $L$-smooth with respect to the negative entropy over the standard simplex. It establishes a fundamental lower bound of $\Omega(L/T)$ on the convergence rate for any first-order optimization algorithm in high dimensions. In contrast to prior work that relied on ill-conditioned constructions, this paper presents the first such lower bound for well-structured proximal settings based on negative entropy, and extends it to the quantum setting involving von Neumann entropy. By integrating tools from convex analysis, relative smoothness theory, complexity lower-bound constructions, and spectral simplex analysis, the work demonstrates that mirror descent is nearly optimal—up to a logarithmic factor—for these problems, and confirms that the same lower bound holds in the quantum regime.
This work proposes a reference-set-free adaptive convergence metric for multi-objective optimization that addresses the scalability limitations of existing indicators when the true Pareto front is unknown. By leveraging the Karush–Kuhn–Tucker (KKT) optimality conditions, the method integrates an entropy-inspired stationarity measure with a quantile normalization mechanism to enhance robustness against heterogeneous residual distributions. While preserving the intrinsic interpretability of KKT-based analysis, the proposed metric significantly improves stability and applicability in both many-objective and high-dimensional scenarios, thereby overcoming the scalability bottlenecks inherent in conventional convergence indicators.
本文解决了光滑$\ell_p/\ell_q$非对偶凸优化问题,通过结合选择器移动与Hölder下降法,提出了一种一阶方法,在高维情况下达到几乎最优的加速效果。