Score
Designs and analyzes optimization algorithms that use proximal operators executed approximately rather than exactly, including proximal gradient and related proximal optimization methods; builds implementations that allow approximate proximal steps and proves conditions for convergence under inexact updates. Develops algorithmic strategies—such as exploiting product structure or efficient transforms—to make proximal steps computationally efficient while preserving theoretical guarantees.
Addressing the lack of a general algorithmic design framework for composite optimization, this paper proposes a systematic transformation framework that converts optimal first-order methods for unconstrained smooth optimization into algorithms for composite optimization. The framework unifies convergence analysis via algebraic identities, eliminating the need for problem-specific redesign or reproof, and naturally generalizes existing methods. It establishes, for the first time, a structural connection between optimal methods for smooth and composite optimization, revealing and exploiting a step-size acceleration mechanism that significantly improves the convergence rate of FISTA-type algorithms—and breaks the current best-known convergence rate bound for gradient norm minimization. Integrating proximal-gradient ideas with computer-assisted algebraic verification, the framework yields several new algorithms. Theoretical analysis proves their superiority over classical methods in both objective-value convergence and gradient-norm decay, demonstrating the framework’s generality and effectiveness.
Accelerating projected/proximal gradient methods for constrained and composite convex optimization—particularly whether convergence can be improved via step-size scheduling alone, without momentum. Method: We propose a novel step-size sequence based on the silver ratio ρ = 1 + √2 and introduce a Laplacian-structured sum-of-squares (SOS) certificate to characterize proximal updates, combined with a recursive concatenation analysis technique. Results: We provide the first rigorous proof that this step-size schedule achieves asymptotically optimal convergence rates: O(ε⁻ˡᵒᵍᵣ²) for smooth convex objectives and O(κˡᵒᵍᵣ² log(1/ε)) for μ-strongly convex objectives with condition number κ, matching the silver-rate lower bound and substantially outperforming fixed or standard decaying step sizes. Our analysis reveals that judicious step-size design—devoid of momentum—is inherently capable of universal acceleration in proximal and projected gradient methods.
This work addresses the weak theoretical convergence guarantees of adaptive optimization algorithms—particularly those employing moving-average momentum estimators—in nonconvex optimization. Methodologically, it establishes a unified convergence analysis framework: (i) it provides the first rigorous proof that monotonic increase of the first-order momentum parameter ensures convergence; (ii) it uncovers a phased co-adaptation mechanism between momentum and step size that enables acceleration; and (iii) it extends the analysis to composite, minimax, and bilevel optimization settings. Theoretical contributions include: (i) nonconvex convergence guarantees applicable to broad classes of adaptive methods (e.g., Adam-type); and (ii) novel, efficient minimax and bilevel optimization algorithms that avoid large batch sizes or double-loop schemes. Empirical results confirm improved convergence rates and generalization performance, corroborating the theoretical insights.
This work addresses stochastic convex composite optimization under a weak noise assumption—namely, only finite variance of stochastic gradients is assumed, without requiring stronger tail conditions (e.g., sub-Gaussianity). We propose a novel stochastic proximal point method that integrates variance reduction with proximal updates and incorporates an efficient iterative solver for the resulting subproblems. Theoretically, under bounded gradient variance, our method achieves high-probability convergence to an $varepsilon$-accurate solution with $O(1/varepsilon)$ sample complexity—strictly improving upon standard SGD-type algorithms. Our key contributions are threefold: (i) eliminating the need for strong noise assumptions prevalent in prior high-probability analyses; (ii) establishing the first low-sample-complexity, high-probability convergence framework for stochastic composite optimization; and (iii) unifying treatment of nonsmooth problem structure and stochastic gradient error within a single algorithmic and analytical framework.
This work addresses the convergence guarantees of stochastic line search optimization for over-parameterized models under interpolation conditions. We establish a necessary and sufficient condition on the search direction—applicable to a broad class of methods—that ensures finite termination and bounded backtracking steps, and rigorously prove linear convergence under the Polyak–Łojasiewicz (PL) assumption. The condition unifies major first-order strategies—including momentum, conjugate gradient, and adaptive preconditioning—providing a verifiable theoretical foundation for their principled integration with stochastic line search. Our analysis fills a critical gap in the convergence theory of stochastic line search methods and significantly extends both the applicability and reliability of efficient first-order optimization in interpolation learning regimes.
This work addresses the slow last-iterate convergence of the Extragradient method in unconstrained bilinear minimax optimization by introducing an acceleration mechanism based on dynamic stepsize scheduling. By formulating stepsize selection as an optimization problem, the paper establishes—for the first time—that convergence rates can be improved using only dynamically adjusted stepsizes, leading to a deterministic stepsize scheme following a power-law decay. The approach is further extended by allowing distinct stepsizes for the extrapolation and update steps, thereby approaching the optimal convergence rate. Under synchronized stepsizes, the method achieves a convergence rate of $O(T^{-2/3 + \varepsilon})$, which is shown to be tight in this setting; with asynchronous stepsizes, the rate improves to nearly optimal $O(T^{-1 + \varepsilon})$.
This work investigates whether gradient descent algorithms relying solely on predetermined non-negative stepsize schedules can achieve the optimal $O(T^{-2})$ last-iterate convergence rate in smooth convex optimization. By constructing adversarial instances and employing a refined recursive analysis, the authors establish—for the first time—a lower bound of $\Omega(T^{-1.9319})$ for this class of methods, rigorously demonstrating that stepsize scheduling alone is insufficient to attain the $O(T^{-2})$ rate. This result delineates the fundamental theoretical limitations of stepsize-scheduled gradient descent and fills a critical gap in the lower-bound analysis for such algorithms.
This study investigates the sensitivity of stochastic optimization algorithms to stepsize choices and the resulting performance degradation. Through theoretical analysis, it establishes a quantitative relationship between stepsize and convergence bounds, providing the first direct theoretical evidence for the robustness of adaptive stepsize methods—such as SPS and NGN—relative to standard SGD, thereby moving beyond prior reliance on empirical comparisons alone. Both theoretical derivations and numerical experiments consistently demonstrate that adaptive methods exhibit significantly greater stability with respect to stepsize selection and more controllable performance degradation. This robustness holds in both convex and non-convex settings, with the latter underscoring the continued theoretical relevance of adaptive strategies even in challenging non-convex optimization landscapes.
This work addresses optimization problems with convex constraints whose intersection is difficult to project onto, covering both strongly convex smooth and general nonsmooth convex settings. The authors propose a novel algorithm that integrates stochastic feasibility methods with (sub)gradient descent, wherein each iteration randomly samples a subset of constraints and employs an adaptive Polyak stepsize that requires no prior knowledge of problem parameters, complemented by iterate averaging. Theoretical analysis establishes linear convergence under strong convexity and a worst-case rate of $O(1/\sqrt{T})$ for general convex objectives, while the infeasibility measure decays geometrically almost surely. Numerical experiments on QCQP and SVM tasks demonstrate superior computational efficiency over existing methods, and under specific sampling strategies, the algorithm achieves optimal convergence rates.
This work addresses the limitation of stochastic proximal algorithms in composite convex optimization, which typically rely on the bounded variance assumption. Focusing on problems with a smooth loss plus a nonsmooth regularizer, the study establishes last-iterate convergence guarantees for both stochastic proximal gradient and stochastic incremental proximal methods. Under the mild assumptions that component functions are convex and smooth—without requiring bounded variance—the authors prove for the first time that both algorithms achieve a near-optimal $\widetilde{O}(1/\sqrt{T})$ convergence rate in the last iterate. This result is directly applicable to practical settings such as graph-guided regularization and extends naturally to structured optimization problems arising in multi-task learning and federated learning.