Score
Designs and implements mathematical and empirical analyses and predictive models of optimization trajectories and optimizer behavior, covering flow and contrastive dynamical analyses, learned (neural) optimizers, and projected-extragradient-style updates. This includes deriving analytic formulas for norm evolution, characterizing scale‑invariant loss effects, proving support‑control and convergence properties, and relating gradients to the time‑course of optimizer state.
This study addresses the theoretical gap in convergence and generalization analyses of deep learning optimization algorithms. We establish a unified theoretical framework encompassing first- and second-order gradient methods, adaptive algorithms (e.g., Adam, K-FAC), and decentralized distributed optimization (e.g., Gossip), providing the first joint analysis of convergence rates and generalization error bounds under non-convex settings. Methodologically, we integrate Lyapunov stability analysis, stochastic optimization convergence proofs, and generalization bound derivation. Key contributions are: (1) filling a critical void in existing surveys by delivering rigorous, self-contained theoretical derivations; (2) pioneering the incorporation of non-convex decentralized optimization into a unified analytical framework; and (3) releasing the first comprehensive theoretical handbook dedicated to deep learning optimization—offering verifiable design principles and advancing understanding of the interplay among training dynamics, solution selection, and generalization.
Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.
This work addresses the reliance on empirical design and poor generalizability of hand-crafted optimizers in gradient-based learning. Methodologically, it formulates the optimizer as a learnable functional mapping from gradients to parameter updates, and—novelty—systematically recasts this as a sequence of analytically solvable convex optimization problems. This unified framework yields closed-form derivations of mainstream optimizers (e.g., SGD, Adam) along with their theoretically optimal hyperparameters. Furthermore, it incorporates a runtime gradient statistics mechanism enabling dynamic, adaptive tuning during training. Experiments demonstrate substantial improvements in convergence speed and training stability, while preserving theoretical rigor and practical deployability.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
There exists a significant gap between the theoretical convergence guarantees of deep learning optimization algorithms and their empirical performance, largely due to commonly adopted assumptions—such as Hessian boundedness—that lack empirical validation. Method: We introduce the first trajectory-aware measurement framework tightly aligned with key theoretical quantities, systematically evaluating the validity of mainstream assumptions across diverse architectures and datasets using large-scale training runs. Our framework quantifies dynamic properties—including gradient norms, Hessian spectral characteristics, and loss curvature—along optimization trajectories. Contribution/Results: We find that all examined theoretical assumptions fail to reliably predict actual convergence behavior and exhibit no robust correlation with optimization performance. This work uncovers a fundamental misalignment between theoretical modeling and practice, establishing the first reproducible benchmark for empirically calibrating and reconstructing optimization theory.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This work addresses the challenge of characterizing the highly complex loss landscape in large language model (LLM) pretraining, where existing theories struggle to balance analytical tractability with accurate dynamic prediction. By performing Taylor expansions of both the model and loss function at mid-training, the authors construct a local quadratic approximation and combine it with Lanczos quadrature and Hessian spectral estimation. For the first time, they validate this approach on a 150M-parameter LLM trained on 3B tokens, demonstrating predictive accuracy over a training window spanning 10% of total steps. Their analysis reveals that the quadratic model faithfully captures optimization trajectories, that the tail structure of the Hessian spectrum is strongly influenced by batch size, preconditioning, and training stage, and that optimization typically resides in a stochastic edge-of-stability regime dictated by batch size—uncovering a deep connection between local stability and hyperparameter choice.
This work addresses the challenge in performative prediction where model deployment induces distributional shifts that complicate optimization. Existing approaches often rely on strong assumptions about the loss function and data distribution, limiting their applicability. To overcome this, the paper proposes a gradient-based adaptive optimization algorithm that explicitly estimates deployment-induced distribution shifts via finite differences, thereby accommodating a broader class of losses and distributions without stringent assumptions. The method supports high-dimensional optimization and incorporates a sample-efficient approximation strategy to reduce data requirements. Theoretical analysis establishes convergence guarantees for the proposed algorithm. Empirical results demonstrate that it converges faster and more stably than existing methods, exhibiting superior robustness and practicality across diverse experimental settings.
This work addresses the lack of a unified framework in existing optimizer design, which often relies on heuristic modifications and struggles to balance stability and generalization. The authors propose the first systematic approach that integrates control theory with Riemannian geometry, modeling the optimization process as a discrete-time controlled dynamical system on a Riemannian manifold. By introducing normally attracting invariant manifolds (NAIMs) and strict Lyapunov functions, they establish a theoretically grounded framework for generating optimizers with provable convergence guarantees. This framework not only recovers classical algorithms but also yields novel optimizers that achieve state-of-the-art performance on large-scale benchmarks. Geometric diagnostics further validate the method’s efficacy, offering a stable, interpretable, and theoretically rigorous toolkit for optimizer design.
This work addresses the high computational cost of evaluating objective functions and their gradients, as well as slow convergence, in engineering optimization. It proposes the Learned Gradient Flow (LGF) optimizer, which employs a data-driven equation discovery approach to infer continuous-time dynamical systems from optimization trajectories—systems that correspond to algorithms such as gradient descent, Newton’s method, and ADAM. LGF constructs surrogate gradient flow models that can replace the original problem, adaptively generating polynomial surrogates of varying orders in either full-dimensional or reduced-dimensional spaces. This significantly reduces reliance on repeated evaluations of the original objective function and its gradients. Demonstrated across diverse forward and inverse problems in structural topology optimization and scientific machine learning, the method accelerates convergence while preserving essential features of the optimization trajectory.
This work addresses the lack of a unified modular framework for analyzing adaptive optimizers, which hinders a precise characterization of their behavior under constraints on directional reachability, information budgets, and update rules. We propose a geometric–non-geometric decoupled calculus for optimizers: the geometric module, constituted by a family of positive-definite cometrics, captures realizable descent directions, while the non-geometric module governs mechanisms such as information processing, memory, and control. Within this framework, we establish a direction expressivity theorem and a residual theory for constrained cometric families, disentangling directional expressiveness from condition-number complexity and recasting optimizer design as a Pareto optimization problem under modular budgets. Theoretically, we prove that fully positive-definite geometry exactly spans all strictly descending directions; experiments demonstrate that high-information full-metric probes attain numerical precision on deterministic quadratic problems, and a Muon-style implementation preliminarily validates the auditability of matrix-operator updates.