analyze optimization dynamics

Designs and implements mathematical and empirical analyses and predictive models of optimization trajectories and optimizer behavior, covering flow and contrastive dynamical analyses, learned (neural) optimizers, and projected-extragradient-style updates. This includes deriving analytic formulas for norm evolution, characterizing scale‑invariant loss effects, proving support‑control and convergence properties, and relating gradients to the time‑course of optimizer state.

analyzeoptimizationdynamics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

From Learning to Optimize to Learning Optimization Algorithms

May 28, 2024
CC
Camille Castera
🏛️ University of Tübingen | Saarland University

Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.

Designing learned optimization algorithms usable beyond training settingsDeveloping learning-enhanced BFGS algorithm adaptable to various test settingsSynergy between classical optimization and Learning to Optimize (L2O)

Optimizing Optimizers for Fast Gradient-Based Learning

Dec 06, 2025
JL
Jaerin Lee
🏛️ Seoul National University

This work addresses the reliance on empirical design and poor generalizability of hand-crafted optimizers in gradient-based learning. Methodologically, it formulates the optimizer as a learnable functional mapping from gradients to parameter updates, and—novelty—systematically recasts this as a sequence of analytically solvable convex optimization problems. This unified framework yields closed-form derivations of mainstream optimizers (e.g., SGD, Adam) along with their theoretically optimal hyperparameters. Furthermore, it incorporates a runtime gradient statistics mechanism enabling dynamic, adaptive tuning during training. Experiments demonstrate substantial improvements in convergence speed and training stability, while preserving theoretical rigor and practical deployability.

Automating optimizer design in gradient-based learningDynamically tuning hyperparameters based on gradient statisticsMaximizing instantaneous loss decrease via optimizer formulation

Data-Driven Performance Guarantees for Classical and Learned Optimizers

Apr 22, 2024
RS
Rajiv Sambharya
🏛️ Princeton University

This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.

Analyzing performance of classical and learned optimization algorithmsDeveloping tighter performance bounds than worst-case guaranteesProviding generalization guarantees using statistical learning theory

Empirical Tests of Optimization Assumptions in Deep Learning

Jul 01, 2024
HT
Hoang Tran
🏛️ Boston University

There exists a significant gap between the theoretical convergence guarantees of deep learning optimization algorithms and their empirical performance, largely due to commonly adopted assumptions—such as Hessian boundedness—that lack empirical validation. Method: We introduce the first trajectory-aware measurement framework tightly aligned with key theoretical quantities, systematically evaluating the validity of mainstream assumptions across diverse architectures and datasets using large-scale training runs. Our framework quantifies dynamic properties—including gradient norms, Hessian spectral characteristics, and loss curvature—along optimization trajectories. Contribution/Results: We find that all examined theoretical assumptions fail to reliably predict actual convergence behavior and exhibit no robust correlation with optimization performance. This work uncovers a fundamental misalignment between theoretical modeling and practice, establishing the first reproducible benchmark for empirically calibrating and reconstructing optimization theory.

Evaluating theoretical optimization analysis methods for deep learningInvestigating practical validity of optimization assumptions and identitiesMeasuring standard analyses' ability to explain modern algorithms

Iterative Linear Quadratic Optimization for Nonlinear Control: Differentiable Programming Algorithmic Templates

Jul 13, 2022
VR
Vincent Roulet
🏛️ Google Brain | University of Washington

This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.

Compare gradient descent, Gauss-Newton, Newton methodsOptimize nonlinear control using differentiable programmingTest algorithms on benchmarks like car racing

Latest Papers

What's happening recently
View more

This work addresses the challenge of characterizing the highly complex loss landscape in large language model (LLM) pretraining, where existing theories struggle to balance analytical tractability with accurate dynamic prediction. By performing Taylor expansions of both the model and loss function at mid-training, the authors construct a local quadratic approximation and combine it with Lanczos quadrature and Hessian spectral estimation. For the first time, they validate this approach on a 150M-parameter LLM trained on 3B tokens, demonstrating predictive accuracy over a training window spanning 10% of total steps. Their analysis reveals that the quadratic model faithfully captures optimization trajectories, that the tail structure of the Hessian spectrum is strongly influenced by batch size, preconditioning, and training stage, and that optimization typically resides in a stochastic edge-of-stability regime dictated by batch size—uncovering a deep connection between local stability and hyperparameter choice.

Hessian spectrumlarge language modelsloss landscape

This work addresses the challenge in performative prediction where model deployment induces distributional shifts that complicate optimization. Existing approaches often rely on strong assumptions about the loss function and data distribution, limiting their applicability. To overcome this, the paper proposes a gradient-based adaptive optimization algorithm that explicitly estimates deployment-induced distribution shifts via finite differences, thereby accommodating a broader class of losses and distributions without stringent assumptions. The method supports high-dimensional optimization and incorporates a sample-efficient approximation strategy to reduce data requirements. Theoretical analysis establishes convergence guarantees for the proposed algorithm. Empirical results demonstrate that it converges faster and more stably than existing methods, exhibiting superior robustness and practicality across diverse experimental settings.

distribution shiftgradient-based methodsloss functions

This work addresses the lack of a unified framework in existing optimizer design, which often relies on heuristic modifications and struggles to balance stability and generalization. The authors propose the first systematic approach that integrates control theory with Riemannian geometry, modeling the optimization process as a discrete-time controlled dynamical system on a Riemannian manifold. By introducing normally attracting invariant manifolds (NAIMs) and strict Lyapunov functions, they establish a theoretically grounded framework for generating optimizers with provable convergence guarantees. This framework not only recovers classical algorithms but also yields novel optimizers that achieve state-of-the-art performance on large-scale benchmarks. Geometric diagnostics further validate the method’s efficacy, offering a stable, interpretable, and theoretically rigorous toolkit for optimizer design.

control theoryconvergenceLyapunov stability

This work addresses the high computational cost of evaluating objective functions and their gradients, as well as slow convergence, in engineering optimization. It proposes the Learned Gradient Flow (LGF) optimizer, which employs a data-driven equation discovery approach to infer continuous-time dynamical systems from optimization trajectories—systems that correspond to algorithms such as gradient descent, Newton’s method, and ADAM. LGF constructs surrogate gradient flow models that can replace the original problem, adaptively generating polynomial surrogates of varying orders in either full-dimensional or reduced-dimensional spaces. This significantly reduces reliance on repeated evaluations of the original objective function and its gradients. Demonstrated across diverse forward and inverse problems in structural topology optimization and scientific machine learning, the method accelerates convergence while preserving essential features of the optimization trajectory.

computational efficiencyequation discoverygradient flow

This work addresses the lack of a unified modular framework for analyzing adaptive optimizers, which hinders a precise characterization of their behavior under constraints on directional reachability, information budgets, and update rules. We propose a geometric–non-geometric decoupled calculus for optimizers: the geometric module, constituted by a family of positive-definite cometrics, captures realizable descent directions, while the non-geometric module governs mechanisms such as information processing, memory, and control. Within this framework, we establish a direction expressivity theorem and a residual theory for constrained cometric families, disentangling directional expressiveness from condition-number complexity and recasting optimizer design as a Pareto optimization problem under modular budgets. Theoretically, we prove that fully positive-definite geometry exactly spans all strictly descending directions; experiments demonstrate that high-information full-metric probes attain numerical precision on deterministic quadratic problems, and a Muon-style implementation preliminarily validates the auditability of matrix-operator updates.

adaptive optimizersdirection expressivitygeometric calculus

Hot Scholars

AO

Antonio Orvieto

ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems
Deep LearningMachine LearningOptimizationDifferential Equations
JY

Jinghui Yuan

Oak Ridge National Laboratory
SafetyCAVSmart MobilityArtificial Intelligence
ZL

Zhouchen Lin

Professor, Peking University; Fellow of IEEE, IAPR, CSIG & AAIA; ex-VP of Samsung Research
machine learningcomputer visionimage processingnumerical optimization
JD

Jason D. Lee

Associate Professor of EECS & Statistics at UC Berkeley
Machine Learning TheoryMachine LearningArtificial IntelligenceStatistics