Score
Designs and carries out mathematical analyses and proofs that characterize the convergence behavior of optimization and learning algorithms, including establishing global and local convergence conditions, asymptotic and finite‑time convergence rates (e.g., o(1/t), o(1/√t)), and statistical/stability guarantees under stochastic or heavy‑tailed noise. Work includes deriving explicit rate‑of‑convergence bounds, bounding optimization and projection error, handling nonconvex and nonsmooth objectives, analyzing adaptive or bias‑corrected methods (e.g., Adam), and using tools such as subspace embeddings and gradient‑dominance to produce explicit convergence ranges and theoretical guarantees.
This work addresses the convergence of RMSProp and Adam for generalized smooth nonconvex optimization. Under the weakest known assumptions—coordinate-wise generalized smoothness and affine noise variance—we establish the first tight theoretical guarantees. We introduce a novel descent lemma that overcomes critical challenges: adaptive step-size dependence, unbounded gradient estimates, and mismatched Lipschitz constants. Rigorously, we prove that both algorithms converge to an ε-stationary point in O(ε⁻⁴) iterations—the optimal rate matching the fundamental lower bound for nonconvex stochastic optimization. Our analysis operates under strictly weaker assumptions and yields tighter bounds than all prior works on RMSProp and Adam, thereby advancing the foundational understanding and reliability certification of adaptive optimization methods.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
There exists a significant gap between the theoretical convergence guarantees of deep learning optimization algorithms and their empirical performance, largely due to commonly adopted assumptions—such as Hessian boundedness—that lack empirical validation. Method: We introduce the first trajectory-aware measurement framework tightly aligned with key theoretical quantities, systematically evaluating the validity of mainstream assumptions across diverse architectures and datasets using large-scale training runs. Our framework quantifies dynamic properties—including gradient norms, Hessian spectral characteristics, and loss curvature—along optimization trajectories. Contribution/Results: We find that all examined theoretical assumptions fail to reliably predict actual convergence behavior and exhibit no robust correlation with optimization performance. This work uncovers a fundamental misalignment between theoretical modeling and practice, establishing the first reproducible benchmark for empirically calibrating and reconstructing optimization theory.
This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.
This paper addresses the theoretical advantages of adaptive methods—specifically AdaGrad—in stochastic nonconvex optimization. Departing from conventional assumptions of global Lipschitz continuity and uniform noise variance, it introduces coordinate-wise fine-grained smoothness and heterogeneous noise variance conditions. Crucially, it adopts the ℓ₁-norm to measure stationarity—i.e., proximity to a gradient stationary point—for the first time in this context. Theoretically, it establishes that AdaGrad achieves an iteration complexity of O(1/ε²), strictly improving upon SGD’s O(d/ε²) and yielding a d-fold acceleration. An information-theoretic lower bound is constructed to confirm that this upper bound is tight up to logarithmic factors. Furthermore, the work develops a novel convergence analysis framework based on the ℓ₁-norm, uncovering the intrinsic advantage of adaptive step sizes in handling coordinate-wise heterogeneity of problem structure and noise.
This work addresses the finite-time convergence of stochastic iterative algorithms for fixed-point equations accessible only through a noisy oracle. The authors propose a norm-independent, unified Lyapunov function framework constructed via a generalized Moreau envelope, which integrates Lyapunov stability theory with stochastic approximation analysis. This framework accommodates complex settings such as Markovian noise, seminorm contractive operators, and dissipative operators, yielding sharp non-asymptotic convergence bounds in both high-probability and mean-square senses. As a result, it provides a unified and refined finite-time convergence guarantee for a broad class of algorithms, including stochastic gradient descent, linear stochastic approximation, Q-learning, and temporal difference learning.
This work addresses the lack of convergence guarantees for the Adam optimizer under heavy-tailed noise, particularly when stochastic gradients possess only bounded centered $p$-th moments with $p \in (1,2]$. We extend the online-to-nonconvex conversion framework to settings involving heavy-tailed martingale difference noise and, by integrating discounted regret analysis with $p$-th moment-constrained optimization techniques, establish the first convergence guarantee for the standard vector-form Adam algorithm. Specifically, we prove that Adam converges to a $(\rho,\varepsilon)$-stationary point, achieving a $p$-dependent suboptimal iteration complexity. Moreover, if the domain radius is known and leveraged to control the outputs of the online learner, the method attains optimal iteration complexity.
This work proposes Asymptotic Learning Theory (ALT), a novel framework designed to efficiently and accurately estimate unknown constants or parameters from known asymptotic forms, while providing rigorous guarantees on convergence and convergence rates. The core methodology introduces sliding Linear Least Squares (sLLSQ) and its Tikhonov-regularized variant (sT-LLSQ), integrating tools from optimization, asymptotic analysis, complex analysis, and analytic combinatorics. For the first time, a theoretical foundation for asymptotic learning is established, proving convergence properties of the proposed estimators and delineating their regimes of applicability. Numerical experiments in contexts such as analytic combinatorics validate the theoretical findings, demonstrating that the method significantly outperforms conventional techniques like the ratio method.
This work addresses the lack of convergence guarantees for classical stochastic gradient methods under heavy-tailed noise—characterized by only finite first- or second-order moments—particularly in unbounded or non-convex settings. The authors propose a unified and concise analytical framework that, without algorithmic modifications or strong assumptions such as bounded domains, establishes for the first time expected convergence of SGD and SGDM in non-convex problems, and of SMD and ASMD in convex problems. By integrating techniques from stochastic mirror descent, momentum acceleration, and moment-condition handling under heavy-tailed distributions, the analysis overcomes key limitations in existing theory and provides a rigorous foundation for stochastic optimization in the presence of heavy-tailed noise.