Score
Designs, implements, and integrates iterative second‑order optimization components such as damped Newton and quasi‑Newton update rules, Hessian‑vector product routines, and compact or normalized Hessian approximations for use inside solvers. Analyzes Hessian structure and spectrum, estimates curvature directions and singular values (including pruning low‑energy directions), and develops query‑efficient procedures to recover or approximate Hessian information from comparisons or limited queries to support preconditioning, uncertainty quantification, and accurate subproblem solving.
This work addresses the efficient and robust approximation of the Hessian (or its inverse) in stochastic optimization, unifying the formulation across Euclidean space, the symmetric positive-definite (SPD) manifold, and general Lie groups. We establish, for the first time, that under mild conditions the Hessian-fitting objective on a Lie group is strongly convex—transforming the original ill-conditioned problem into a well-conditioned Lie-group optimization task. Leveraging this property, we propose an efficient preconditioner construction method based on Hessian-vector products and sparse structural priors. We further conduct a systematic geometric analysis comparing the adaptability and convergence behavior of second-order and adaptive methods—including BFGS, natural gradient, and PSGD—within this unified framework. Theoretical analysis and empirical evaluation demonstrate that our approach significantly improves the efficiency, numerical stability, and scalability of second-order information utilization in large-scale stochastic optimization.
This work addresses the longstanding absence of second-order derivatives (Hessians) for implicit functions in differentiable finite element physics. We present the first systematic derivation and implementation of an implicit Hessian-vector product algorithm based on primitive automatic differentiation operators, enabling PDE-constrained optimization. Our method leverages Jacobian-vector and vector-Jacobian products as core primitives, enabling seamless integration of Newton-CG and L-BFGS-B optimizers into finite element solvers. We establish a verifiable and scalable second-order implicit differentiation paradigm, thereby bridging a critical theoretical and practical gap in differentiable physics concerning Hessian computation. The approach is validated on four 2D/3D benchmark problems—spanning both linear and nonlinear regimes—with verified numerical accuracy. When combined with exact Hessians, Newton-CG accelerates convergence by 2–5× for traction identification and shape optimization tasks.
This work addresses the sensitivity to initial guesses and high computational cost of Newton’s method for solving nonlinear parameterized partial differential equations. The authors propose a two-stage initialization strategy: first, by leveraging parameter sampling and a precomputed solution library, they construct two complementary feature spaces—solution manifold and corrected search directions—from discrete Newton trajectories; second, a regression model predicts a surrogate initial guess, which is then refined via lightweight GMRES-based residual minimization to yield a high-quality starting point. Operating under a weakly intrusive framework, this approach significantly accelerates high-fidelity Newton iterations, markedly reducing both iteration counts and total CPU time on benchmark PDE problems, outperforming existing methods that rely solely on surrogate-based initialization.
This work addresses unconstrained nonconvex optimization under Hölder-continuous Hessians, aiming to efficiently compute approximate first- and second-order stationary points. We propose a novel Newton-CG method integrating adaptive regularization with a dynamic conjugate gradient (CG) stopping criterion—requiring no prior knowledge of the Hölder exponent. It is the first fully parameter-free second-order algorithm achieving optimal theoretical complexity: both iteration count and Hessian-vector product count match established lower bounds. The method also supports explicit parameter-dependent variants for enhanced flexibility. We establish theoretical guarantees for simultaneous convergence to an ε-first-order stationary point and an ε^{1/2}-second-order stable point. Numerical experiments demonstrate superior convergence speed, robustness, and reduced hyperparameter sensitivity compared to classical regularized Newton methods.
This work addresses the prohibitively high computational cost of second-order methods in convex–concave minimax optimization. We propose a “lazy Hessian” strategy that maintains optimal iteration complexity $mathcal{O}(varepsilon^{-3/2})$ while drastically reducing per-iteration cost. By reusing Hessian information across iterations and incorporating an adaptive update mechanism, we achieve the first total computational complexity of $ ilde{mathcal{O}}ig((N + d^2)(d + d^{2/3}varepsilon^{-2/3})ig)$, improving upon the Monteiro–Svaiter (2012) optimal method by a factor of $d^{1/3}$. We further extend the framework to the strongly convex–strongly concave setting. The theoretical analysis is rigorous, and extensive numerical experiments on both synthetic and real-world datasets confirm accelerated convergence and reduced resource consumption.
This work addresses the limited convergence rate of the BFGS algorithm in quasi-Newton methods by proposing a novel search direction termed the quasi-quadratic gradient (QQG). The QQG method explicitly incorporates local second-order curvature information into the construction of the search direction for the first time, achieved by dynamically correcting the optimization trajectory through multiplication of the inverse Hessian approximation with the current gradient. While preserving the computational efficiency inherent to BFGS, this approach significantly accelerates convergence. Numerical experiments demonstrate that QQG achieves faster convergence rates than standard BFGS across a variety of benchmark problems, thereby validating the effectiveness and superiority of explicitly leveraging local curvature information in the design of search directions.
This work addresses the numerical instability and convergence failure of L-BFGS in ill-conditioned or nonconvex optimization problems, which arise when the condition number of the inverse Hessian approximation grows uncontrollably. To mitigate this issue, the authors propose Two-Sided L-BFGS, a novel variant that incorporates a bilateral geometric envelope mechanism to dynamically bound the condition number while preserving the standard computational complexity and curvature information. The method establishes, for the first time, an explicit upper bound on the condition number of the L-BFGS inverse Hessian approximation, revealing its dependence on memory depth, problem dimensionality, and envelope hyperparameters. Global convergence is guaranteed even in nonconvex settings. Experimental results demonstrate that the proposed approach significantly enhances robustness and convergence performance on high-dimensional ill-conditioned problems.
This work addresses the lack of finite-time convergence guarantees in zeroth-order multi-timescale stochastic optimization by studying two-timescale gradient and three-timescale Newton methods that rely solely on function-value feedback. By employing smoothing functionals to estimate gradients and Hessians, the paper establishes the first non-asymptotic convergence analysis for zeroth-order multi-timescale algorithms, explicitly characterizing the coupling between timescales and the associated error propagation mechanisms. Key contributions include deriving mean-squared error bounds for Hessian estimation, providing a finite-time upper bound on the norm of the objective gradient, proving convergence to a first-order stationary point, and proposing a stepsize strategy that balances dominant error sources to achieve a near-optimal convergence rate. The theoretical findings are validated in the Continuous Mountain Car environment.