Score
Design, implement, and analyze iterative numerical solvers and their preconditioners for linear and nonlinear systems, including development of preconditioned iterative methods and the design of effective preconditioners. Specify and validate convergence and stopping criteria, integrate solvers with learned models and user-driven iteration workflows, and devise strategies to estimate solutions for underdetermined systems.
Existing preconditioners for large-scale sparse linear systems suffer from low efficiency and poor generalization. Method: This paper proposes a novel learnable preconditioner that integrates algebraic preconditioning with graph neural networks (GNNs). It initializes the GNN with a classical ILU-type preconditioner and introduces a differentiable, condition-number-based loss function to explicitly optimize spectral properties during training. Additionally, it incorporates sparse structural priors and parameterized PDE modeling to ensure physical consistency and computational tractability. Results: On benchmark discretized parametric PDE systems, the method reduces iterative solver iterations by 30–50% compared to ILU and state-of-the-art neural preconditioners, achieves significantly improved condition numbers, and incurs only modest inference overhead. This work overcomes key limitations of purely data-driven and purely sparse-GNN-based preconditioners, establishing a new paradigm for interpretable, efficient, and generalizable learning in numerical linear algebra.
This work addresses the convergence guarantees of stochastic line search optimization for over-parameterized models under interpolation conditions. We establish a necessary and sufficient condition on the search direction—applicable to a broad class of methods—that ensures finite termination and bounded backtracking steps, and rigorously prove linear convergence under the Polyak–Łojasiewicz (PL) assumption. The condition unifies major first-order strategies—including momentum, conjugate gradient, and adaptive preconditioning—providing a verifiable theoretical foundation for their principled integration with stochastic line search. Our analysis fills a critical gap in the convergence theory of stochastic line search methods and significantly extends both the applicability and reliability of efficient first-order optimization in interpolation learning regimes.
Traditional algebraic preconditioners (e.g., ILU, AMG) suffer from failure on ill-conditioned large-scale sparse linear systems, exhibit unpredictable and expensive setup costs, and often rely on problem-specific physical priors. Method: We propose the first end-to-end differentiable, general-purpose preconditioner based on graph neural networks (GNNs). It encodes sparse matrices as graphs without requiring underlying physical knowledge, enabling strong generalization across diverse problem domains. Contribution/Results: The GNN-based preconditioner achieves highly predictable and significantly accelerated setup times compared to ILU and AMG. Integrated tightly with Krylov subspace methods (e.g., GMRES), it reduces iteration counts relative to inner-outer GMRES. Evaluated on 800+ real-world matrices spanning PDEs, economics, statistics, and graph learning, our approach consistently improves both solver efficiency and robustness—demonstrating superior scalability, generality, and practical applicability for large-scale sparse linear systems.
Existing iterative solvers for large-scale linear systems suffer from strong dependence on the global condition number and coarse-grained complexity analyses. Method: We introduce the *spectral tail condition number* $kappa_ell$, a new fine-grained spectral measure, and develop a refined time-complexity framework. Our approach formally defines $kappa_ell$, integrates it with the Sketch-and-Project paradigm, Nesterov acceleration, determinant point process sampling, and universality theory for Gaussian matrices, thereby exposing an intrinsic connection between iteration complexity and the matrix multiplication exponent $omega$. Contribution/Results: Our analysis achieves a sharper separation between deterministic and randomized algorithms, yielding an $ ilde{O}(kappa_ell n^2 log(1/varepsilon))$ bound for computing an $varepsilon$-accurate solution—valid for $ell$ up to $O(n^{0.729})$. This significantly improves the fine-grained analysis of the conjugate gradient method and establishes a novel theoretical benchmark for iterative algorithm design.
In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.
This work addresses the high sensitivity of Krylov iterative solvers to geometry, boundary conditions, and material parameters when solving parametric partial differential equations, as well as the limited generalization and acceleration capabilities of existing neural operator-based preconditioners. The authors propose NSPOD, a multigrid-like deep operator network preconditioner that, for the first time, integrates neural operators with Proper Orthogonal Decomposition (POD) subspaces. By approximating solutions within a low-dimensional POD subspace, NSPOD effectively accelerates Krylov solvers without requiring retraining, even on unstructured meshes derived from complex CAD geometries. Demonstrated on linear PDEs in solid mechanics, the method significantly reduces iteration counts and outperforms state-of-the-art preconditioners such as algebraic multigrid, thereby overcoming key performance bottlenecks in current approaches.
This work addresses the limitations of classical iterative methods, which rely on forward error and are constrained by the condition number of the matrix. It introduces a new paradigm using backward error as the convergence criterion. The key contributions include the first proof that Richardson iteration achieves a condition-number-independent $O(1/k)$ convergence rate in backward error for any positive semidefinite linear system. Building on this, the authors design an accelerated algorithm, MINBERR, attaining an $O(1/k^2)$ convergence rate. They further integrate backward error minimization into Krylov subspace methods and extend the approach to general linear systems. The resulting general-purpose solver has complexity $O(n^2/\varepsilon)$, while MINBERR achieves $O(n^2/\sqrt{\varepsilon})$, demonstrating superior numerical performance in benchmark experiments.
This work addresses the sensitivity to initial guesses and high computational cost of Newton’s method for solving nonlinear parameterized partial differential equations. The authors propose a two-stage initialization strategy: first, by leveraging parameter sampling and a precomputed solution library, they construct two complementary feature spaces—solution manifold and corrected search directions—from discrete Newton trajectories; second, a regression model predicts a surrogate initial guess, which is then refined via lightweight GMRES-based residual minimization to yield a high-quality starting point. Operating under a weakly intrusive framework, this approach significantly accelerates high-fidelity Newton iterations, markedly reducing both iteration counts and total CPU time on benchmark PDE problems, outperforming existing methods that rely solely on surrogate-based initialization.
To address the slow convergence and poor generalization of iterative solvers for parametric partial differential equations (PDEs) on arbitrary unstructured meshes, this paper proposes a geometry-aware hybrid preconditioning framework. The method integrates finite-element mesh encoding with a novel Geo-DeepONet architecture—yielding the first neural preconditioner capable of cross-geometry generalization without retraining. Coupled with Krylov subspace methods (e.g., GMRES) and multilevel relaxation strategies, it enables geometry-adaptive iterative acceleration. Evaluated on parametric PDEs from elasticity and heat conduction, the framework reduces generalization error on unseen geometries by over 60%, accelerates overall solution time by 3–5×, and significantly improves robustness and transferability across diverse geometric domains.