curvature diagnostics

Design and implement methods to estimate and track local second-order curvature of objective functions—primarily Hessian eigenvalues and eigenvectors—using efficient Hessian-vector products, stochastic or warm-started power iteration, and preconditioning-aware techniques. Use these online curvature estimates as diagnostics during optimization to report largest preconditioned eigenvalues, detect influential observations under perturbations, compute curvature-based influence measures, and inform preconditioner or step-selection decisions.

curvaturediagnostics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Second-order optimization methods—e.g., Newton’s method—systematically fail in deep neural network training when applied with the exact Hessian, despite its theoretical role in capturing local curvature. Method: We combine nonlinear discrete-time dynamical modeling with exact Hessian analysis of regression loss landscapes to characterize optimization dynamics under discrete parameter updates. Contribution/Results: We reveal a geometric mismatch mechanism: critical points are not densely clustered near local minima but instead form a sparse, high-dimensional saddle-dominated structure. Consequently, exact curvature information induces directional misalignment and step-size instability—not convergence acceleration—during discrete updates. This challenges the long-standing “ubiquitous local minima” hypothesis and provides the first curvature-utilization mismatch explanation for second-order failure in deep learning. Our framework delivers both a novel theoretical perspective on optimization geometry and empirical validation, advancing the understanding of why canonical second-order methods underperform in modern neural network training.

Analyzing geometry of nonlinear discretizations in optimization landscapesChallenging conventional wisdom about local minima distributionInvestigating failure of neural network training with exact Hessian methods

This work addresses the efficient and robust approximation of the Hessian (or its inverse) in stochastic optimization, unifying the formulation across Euclidean space, the symmetric positive-definite (SPD) manifold, and general Lie groups. We establish, for the first time, that under mild conditions the Hessian-fitting objective on a Lie group is strongly convex—transforming the original ill-conditioned problem into a well-conditioned Lie-group optimization task. Leveraging this property, we propose an efficient preconditioner construction method based on Hessian-vector products and sparse structural priors. We further conduct a systematic geometric analysis comparing the adaptability and convergence behavior of second-order and adaptive methods—including BFGS, natural gradient, and PSGD—within this unified framework. Theoretical analysis and empirical evaluation demonstrate that our approach significantly improves the efficiency, numerical stability, and scalability of second-order information utilization in large-scale stochastic optimization.

Analyzes efficiency of preconditioner methods in various geometric settingsInvestigates Hessian fitting for stochastic optimizations using PSGD criterionProves Hessian fitting is strongly convex in certain Lie groups

A Regularized Newton Method for Nonconvex Optimization with Global and Local Complexity Guarantees

Feb 07, 2025
YZ
Yuhao Zhou
🏛️ Tsinghua University | The Hong Kong Polytechnic University | Chinese Academy of Sciences

This work addresses efficient computation of ε-stationary points in nonconvex optimization problems with Lipschitz-continuous Hessians. We propose an adaptive quadratic regularization Newton method: it introduces, for the first time, a gradient-difference-driven adaptive regularization term that eliminates the need to know the Hessian Lipschitz constant a priori; subproblems are solved via a negative-curvature-aware linear conjugate gradient method. Theoretically, the algorithm achieves a global second-order oracle complexity of O(ε⁻³⁄²) and an Hessian-vector product complexity of Õ(ε⁻⁷⁄⁴). Moreover, it attains local quadratic convergence when restricted to regions where the Hessian is positive definite. Numerical experiments demonstrate superior performance over state-of-the-art nonconvex optimizers.

Achieves efficient second-order oracle complexityAdaptive quadratic regularized Newton methodNonconvex optimization with global guarantees

Optimal Preconditioning is a Geodesically Convex Optimization Problem

Dec 06, 2025
ML
M. Levent Doğan
🏛️ Ruhr University Bochum | The University of Texas at San Antonio | INRIA Paris | Sorbonne Université

This work addresses the construction of approximately optimal preconditioners for linear and nonlinear systems to minimize condition numbers. We propose the first unified framework that models condition-number minimization under structured transformations as a geodesically convex optimization problem over unitarily invariant norms. We establish, for the first time, the geodesic convexity of the local condition number of polynomial systems under symmetric Lie group actions, enabling the design of the first preconditioning algorithm with provable convergence guarantees. Leveraging Riemannian gradient updates and Lie group actions, we develop efficient first-order algorithms for Frobenius condition-number optimization, achieving convergence rates of $widetilde{O}(1/varepsilon^2)$ and $widetilde{O}(kappa_F^2 log(1/varepsilon))$, respectively. Both theoretical analysis and empirical evaluation validate the efficacy of our approach.

Extends framework to polynomial systems with theoretical guarantees.Minimizes condition number using geodesically convex optimization methods.Optimizes preconditioners for linear and nonlinear equation systems.

This work addresses unconstrained nonconvex optimization under Hölder-continuous Hessians, aiming to efficiently compute approximate first- and second-order stationary points. We propose a novel Newton-CG method integrating adaptive regularization with a dynamic conjugate gradient (CG) stopping criterion—requiring no prior knowledge of the Hölder exponent. It is the first fully parameter-free second-order algorithm achieving optimal theoretical complexity: both iteration count and Hessian-vector product count match established lower bounds. The method also supports explicit parameter-dependent variants for enhanced flexibility. We establish theoretical guarantees for simultaneous convergence to an ε-first-order stationary point and an ε^{1/2}-second-order stable point. Numerical experiments demonstrate superior convergence speed, robustness, and reduced hyperparameter sensitivity compared to classical regularized Newton methods.

Achieve parameter-free second-order stationary point findingDevelop Newton-CG for nonconvex unconstrained optimizationImprove iteration and operation complexity efficiency

Latest Papers

What's happening recently
View more

This work aims to unify and extend the applicability of Newton-type optimization algorithms by incorporating curvature information into gradient updates in a more general manner. To this end, the authors propose the Generalized Quadratic Gradient (GQG) framework, which abstracts the common structural properties of existing methods and expresses the update rule as the integration of the gradient with any positive-definite curvature matrix satisfying the stationarity condition of a local quadratic model. This framework transcends the conventional reliance on specific Hessian approximations—such as diagonal matrices or BFGS—and establishes a universal optimization paradigm applicable to arbitrary positive-definite curvature matrices. Grounded in a generalized analysis of local quadratic models and quasi-Newton theory, this study provides a rigorous theoretical foundation and methodological guidance for designing more flexible and efficient curvature-aware optimization algorithms.

Hessian approximationNewton-type optimizationoptimization framework

This study addresses the open question of whether high-precision Gauss–Newton curvature and indefinite Hessians can uniformly coexist within low-cost sets under ridge regularization in nonlinear least squares. By integrating local regularity analysis, level-set curvature persistence theory, and structural analysis of Jacobian row rank, this work constructs a pointwise certificate mechanism to characterize error bounds. It provides the first proof that high-accuracy curvature and indefinite Hessians coexist uniformly within such low-cost sets. Furthermore, it establishes the applicability of a single positive ridge upper bound over fixed neighborhoods that does not shrink with weight decay. A theoretical upper bound of (1+√2)/8 on the relative Hessian error is derived, while an indefinite counterexample yielding an error of at least 15/8 is constructed to verify the tightness of this bound.

curvature accuracyGauss-Newton approximationindefinite Hessian

This work addresses the quantification of geometric complexity in optimization problems under prescribed metric structures and provides verifiable certificates for structured preconditioning. By interpreting optimizer states as evolving geometries, the paper introduces a novel framework of “constrained geometric complexity,” which integrates monotonicity principles with submanifold distance theory to derive an exact complexity formula for the two-dimensional diagonal case and a Kronecker projection theorem. The approach combines linear matrix inequalities, Loewner-order squeezing, Hessian-relative certificates, and low-rank spectral models to establish reachability criteria for both diagonal and block structures. It constructs computable mismatch certificates that verify the convergence of Armijo-type solvers and demonstrates empirical efficacy on small-scale quadratic problems.

certificatescondition numbergeometric complexity

This work addresses nonconvex equality-constrained optimization by proposing a gradient-eigenstep algorithm based on the Fletcher augmented Lagrangian function to efficiently compute approximate second-order stationary points. Under suitable initialization and parameter conditions, the algorithm is shown for the first time to enjoy local linear convergence in a neighborhood of strong second-order stationary points. Furthermore, when embedded as a subproblem solver within an incremental sampling strategy, the method significantly outperforms approaches that directly solve the full-sample problem for large-scale stochastic constrained optimization, thereby substantially reducing worst-case sample complexity.

augmented Lagrangianequality-constrained optimizationnonconvex optimization

Hot Scholars

RI

Rustem Islamov

PhD student, University of Basel
machine learningoptimization
HC

Huanran Chen

PhD student, Tsinghua SAIL
Machine Learning TheoryOptimizationAI Safety
JP

Junbiao Pang

Beijing University of Technology
computer visionmultimediamachine learning