iteration-level nonlinear preconditioning

Designs, implements, or analyzes nonlinear preconditioners that are applied at each iteration of a nonlinear solver (iteration-level), specifically right-preconditioning the nonlinear residual or update; this includes constructing iteration operators or learned correction maps that transform residuals or iterates—optionally restricted to active interface variables—to accelerate convergence and prevent stagnation.

iteration-levelnonlinearpreconditioning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address stagnation or divergence of Newton’s method in solving parametric nonlinear systems caused by severe nonlinearity-induced imbalance, this paper proposes a neural operator-preconditioned Newton method. Our approach integrates deep learning priors into classical numerical optimization: (1) we design a fixed-point neural operator (FPNO) that learns, in an end-to-end differentiable manner, the mapping from the current iterate to the true solution—serving as a learned preconditioner; and (2) we introduce an adaptive negative-step mechanism that relaxes reliance on local convexity assumptions inherent in conventional line search and trust-region methods. The method significantly enhances convergence robustness and computational efficiency for strongly nonlinear, multiscale parametric problems. Experiments across multiple real-world physical modeling tasks demonstrate 2–5× speedup over classical solvers, with markedly reduced sensitivity to initial guesses.

Adaptively employing negative step sizes for strong nonlinearitiesOvercoming Newton iteration stagnation using neural operatorsSolving parametric nonlinear systems with unbalanced nonlinearities

Existing preconditioners for large-scale sparse linear systems suffer from low efficiency and poor generalization. Method: This paper proposes a novel learnable preconditioner that integrates algebraic preconditioning with graph neural networks (GNNs). It initializes the GNN with a classical ILU-type preconditioner and introduces a differentiable, condition-number-based loss function to explicitly optimize spectral properties during training. Additionally, it incorporates sparse structural priors and parameterized PDE modeling to ensure physical consistency and computational tractability. Results: On benchmark discretized parametric PDE systems, the method reduces iterative solver iterations by 30–50% compared to ILU and state-of-the-art neural preconditioners, achieves significantly improved condition numbers, and incurs only modest inference overhead. This work overcomes key limitations of purely data-driven and purely sparse-GNN-based preconditioners, establishing a new paradigm for interpretable, efficient, and generalizable learning in numerical linear algebra.

Conjugate Gradient SolverLarge-scale Mathematical ProblemsPreconditioner Efficiency

This work addresses the learning of solution operators for nonlinear partial differential equations (PDEs). Methodologically, it introduces a multilayer perceptron (MLP)-based parametric operator learning framework driven by physical inputs—including boundary conditions, coefficients, and source terms—where a local energy functional, constructed via finite element discretization, is directly adopted as the training loss—a novel formulation. To enhance scalability, the method incorporates element-wise parallelization and randomized sparse grid sampling, significantly improving training efficiency for large-scale problems. Experiments on multiple nonlinear PDE benchmarks demonstrate high-accuracy solution prediction and strong generalization across unseen parameter configurations. Compared to conventional numerical solvers, the approach drastically reduces computational overhead associated with repeated parameter sweeps and enables real-time parametric response prediction.

Combining MLPs with finite element methods for PDE solutionsDeveloping efficient training algorithms via localized energy assemblyLearning solution operators for nonlinear PDEs using MLPs

This work addresses the convergence guarantees of stochastic line search optimization for over-parameterized models under interpolation conditions. We establish a necessary and sufficient condition on the search direction—applicable to a broad class of methods—that ensures finite termination and bounded backtracking steps, and rigorously prove linear convergence under the Polyak–Łojasiewicz (PL) assumption. The condition unifies major first-order strategies—including momentum, conjugate gradient, and adaptive preconditioning—providing a verifiable theoretical foundation for their principled integration with stochastic line search. Our analysis fills a critical gap in the convergence theory of stochastic line search methods and significantly extends both the applicability and reliability of efficient first-order optimization in interpolation learning regimes.

Analyzing convergence of stochastic line search for over-parametrized modelsDefining conditions for finite termination in backtracking proceduresIdentifying fast convergence properties for PL functions in interpolation

Traditional algebraic preconditioners (e.g., ILU, AMG) suffer from failure on ill-conditioned large-scale sparse linear systems, exhibit unpredictable and expensive setup costs, and often rely on problem-specific physical priors. Method: We propose the first end-to-end differentiable, general-purpose preconditioner based on graph neural networks (GNNs). It encodes sparse matrices as graphs without requiring underlying physical knowledge, enabling strong generalization across diverse problem domains. Contribution/Results: The GNN-based preconditioner achieves highly predictable and significantly accelerated setup times compared to ILU and AMG. Integrated tightly with Krylov subspace methods (e.g., GMRES), it reduces iteration counts relative to inner-outer GMRES. Evaluated on 800+ real-world matrices spanning PDEs, economics, statistics, and graph learning, our approach consistently improves both solver efficiency and robustness—demonstrating superior scalability, generality, and practical applicability for large-scale sparse linear systems.

Iterative MethodsPreconditioningSparse Linear Systems

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional iterative methods for solving large-scale systems of equations, including low efficiency, poor robustness, and difficulties in parameter selection. To overcome these challenges, it proposes a "hybrid iteration" paradigm that integrates classical numerical algorithms with machine learning. Building upon conventional iterative schemes such as Newton's method, this approach incorporates deep learning-based optimization strategies to enable adaptive hyperparameter tuning, thereby combining the flexibility of data-driven methods with the theoretical reliability and interpretability of traditional algorithms. Furthermore, this project systematically reviews state-of-the-art methodologies in this domain and identifies key open challenges. Ultimately, it establishes a clear research trajectory for developing efficient and robust solvers for scientific computing.

iterative methodslinear systemsmachine learning

This work addresses the sensitivity to initial guesses and high computational cost of Newton’s method for solving nonlinear parameterized partial differential equations. The authors propose a two-stage initialization strategy: first, by leveraging parameter sampling and a precomputed solution library, they construct two complementary feature spaces—solution manifold and corrected search directions—from discrete Newton trajectories; second, a regression model predicts a surrogate initial guess, which is then refined via lightweight GMRES-based residual minimization to yield a high-quality starting point. Operating under a weakly intrusive framework, this approach significantly accelerates high-fidelity Newton iterations, markedly reducing both iteration counts and total CPU time on benchmark PDE problems, outperforming existing methods that rely solely on surrogate-based initialization.

computational accelerationinitial guessNewton's method

This work addresses the high sensitivity of Krylov iterative solvers to geometry, boundary conditions, and material parameters when solving parametric partial differential equations, as well as the limited generalization and acceleration capabilities of existing neural operator-based preconditioners. The authors propose NSPOD, a multigrid-like deep operator network preconditioner that, for the first time, integrates neural operators with Proper Orthogonal Decomposition (POD) subspaces. By approximating solutions within a low-dimensional POD subspace, NSPOD effectively accelerates Krylov solvers without requiring retraining, even on unstructured meshes derived from complex CAD geometries. Demonstrated on linear PDEs in solid mechanics, the method significantly reduces iteration counts and outperforms state-of-the-art preconditioners such as algebraic multigrid, thereby overcoming key performance bottlenecks in current approaches.

convergence accelerationKrylov-based iterative solversparametric PDEs

This study addresses the ill-conditioning and insufficient accuracy of loss functions induced by differential operators during the training of physics-informed neural operators. To overcome this limitation, this work proposes a preconditioned residual loss that integrates geometric and algebraic multigrid techniques to achieve mesh-independent condition number control. Notably, this strategy is architecture-agnostic and incurs zero overhead during inference. By effectively resolving these optimization challenges, the proposed approach breaks through the bottlenecks of conventional unsupervised physics-informed learning. Experimental results demonstrate that the method attains supervised-level accuracy on benchmark equations such as the Poisson equation, yielding a four- to twenty-five-fold improvement in accuracy over existing state-of-the-art approaches.

Differential operatorsIll-conditioningNeural operators

This work addresses the limited generalization of neural PDE solvers under varying boundary conditions, which arises because these models learn a family of operators conditioned on the training boundary distribution rather than a universal operator. The authors formulate operator learning as a conditional risk minimization problem with respect to boundary conditions and introduce the concept of a “boundary-indexed operator family.” They theoretically prove that standard neural operators suffer from non-identifiability outside the training boundary distribution, revealing the root cause of their generalization bottleneck. Through controlled experiments, numerical simulations of Poisson’s equation, and boundary perturbation analyses, they demonstrate significant performance degradation under boundary shifts and show that removing boundary information reduces the model to a conditional expectation. This study is the first to systematically elucidate the critical role of boundary conditions in the generalization of neural operators.

boundary conditionsgeneralizationneural PDE solvers

Hot Scholars

AS

Antonio Silveti-Falls

CentraleSupélec, Université Paris-Saclay
Nonsmooth OptimizationStochastic OptimizationDeep Learning
VC

Volkan Cevher

Associate Professor, LIONS, EPFL. Amazon Scholar (AGI Foundations).
Machine LearningOptimizationSignal ProcessingInformation Theory
AS

Abdurakhmon Sadiev

PhD student, KAUST
OptimizationFederated LearningVariational Inequalities
RK

Rolf Krause

Full Professor, KAUST
Numerical Solution of PDEsMachine LearningMultigrid/Domain DecompositionContact Problems