gpu algebraic multigrid

Designs, implements, and analyzes algebraic multigrid (AMG) methods and preconditioners for sparse linear systems, including GPU-accelerated AMG implementations and their integration as primitives or wrappers into solver stacks. Builds and tunes configuration, deployment, and coupling of AMG with iterative Krylov solvers on GPUs to optimize sparse solves and preconditioning performance.

gpualgebraicmultigrid

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical gap in the JAX ecosystem: the absence of a sparse linear solver that supports GPU-accelerated algebraic multigrid (AMG), automatic differentiation, and distributed multi-GPU execution. The authors present the first native integration of NVIDIA AmgX as a JAX primitive, unifying AMG with Krylov subspace methods within JAX’s computational framework. This integration enables just-in-time (JIT) compilation, reverse-mode automatic differentiation, batching, and MPI-based distributed execution across multiple GPUs. To enhance efficiency, the implementation incorporates a solver caching mechanism that mitigates repeated setup overhead. By delivering a high-performance, scalable sparse linear algebra layer, this work bridges a key tooling gap between differentiable simulation and scientific computing, enabling seamless incorporation into scientific machine learning pipelines—particularly for PDE-constrained optimization and inverse problems.

algebraic multigridautomatic differentiationGPU acceleration

This work addresses the high data movement overhead and lack of sparse library support in block algebraic multigrid (AMG) due to the nonstandard block structure of the Galerkin triple product $A_c = P^T A P$. The authors propose shared-memory block kernels and a search-free scheduling strategy, along with a novel prolongation operator filtering algorithm based on the Frobenius norm, which significantly reduces data transfer, coarse-operator fill-in, and memory footprint while preserving iterative convergence. A fully GPU-resident block AMG pipeline is implemented using Kokkos (with CUDA backend) and native CUDA, eliminating scalar expansion and host-device data copies. On an NVIDIA A100, DRAM traffic for the fine-grid Galerkin product drops from 17.4 GB to 10.5 GB, execution time decreases from 82 ms to 45 ms, and hotspot operations achieve a 2.9× speedup, approaching theoretical bandwidth limits.

algebraic multigridblock sparse matricesdata movement

This work addresses the inefficiency of existing GPU-based algebraic multigrid (AMG) implementations that rely on scalar sparse matrix formats, which fail to exploit the natural rectangular block structure arising from discretizations of vector-valued PDEs such as elasticity, leading to high memory overhead and low arithmetic intensity. The paper presents the first fully GPU-resident, native block AMG pathway in PETSc, built upon Kokkos to support a portable block-sparse matrix type with heterogeneous block sizes across coarse and fine grids while preserving block structure throughout the smoothed aggregation AMG process. It integrates PetscSF for block prolongation communication across processes and introduces a dense rectangular block COO assembly mechanism. Experiments on 3D elasticity problems using A100 GPUs demonstrate significantly reduced GPU memory usage compared to cuSPARSE and scalar Kokkos approaches, achieving a 2.27× speedup in Galerkin construction at 64-GPU scale, along with 1.24× and 1.42× improvements in V-cycle and SpMV performance, respectively.

algebraic multigridblocked matriceselasticity

Extending DD-$alpha$AMG on heterogeneous machines

Jul 10, 2024
LH
Lianhua He
🏛️ Computer Network Information Center, Chinese Academy of Sciences | Department of Mathematics, Bergische Universität Wuppertal

This work addresses the key challenge of enabling efficient, cross-architecture execution of the DD-αAMG multigrid solver for lattice QCD simulations on heterogeneous exascale supercomputers (e.g., ORISE). Method: We present the first full port of DD-αAMG to the HIP platform, unifying acceleration across both NVIDIA and AMD GPUs; introduce an improved Richardson smoother that achieves superior convergence–speed trade-offs—outperforming GCR significantly and lagging SAP by only ~10%; and extend the odd-even preconditioned GMRES/Richardson hybrid iteration, applying advanced SIMD vectorization and computational restructuring to coarse-grid operations. Contribution/Results: Experiments on ORISE demonstrate efficient heterogeneous acceleration, with substantial performance gains in critical coarse-grained operations. The proposed framework establishes a scalable, cross-platform AMG solver paradigm tailored for large-scale lattice QCD computations.

Extend DD-αAMG for heterogeneous machinesImprove smoothers like Richardson in DD-αAMGPort coarse-grid operations via advanced vectorization

Preconditioning Techniques for Hybridizable Discontinuous Galerkin Discretizations on GPU Architectures

Dec 15, 2025
AW
Andrew Welter
🏛️ Massachusetts Institute of Technology

High-order discontinuous Galerkin (HDG) methods face severe iterative inefficiency and memory bottlenecks when solving PDEs on GPUs, primarily due to the unstructured sparsity and large size of the global stiffness matrix. Method: This work introduces the first GPU-native, block-structured storage format tailored for HDG global stiffness matrices, coupled with a full-stack parallel iterative solver and preconditioning framework. Key techniques include: (1) elimination of local degrees of freedom to maximize arithmetic intensity; (2) batched block-Jacobi, additive Schwarz, and polynomial smoothers, co-designed with shared-memory optimization and memory-coalesced access patterns; and (3) native support for Newton–GMRES nonlinear solvers and domain-decomposition preconditioning. Results: The framework demonstrates high efficiency across six PDE classes—from Poisson to Reynolds-Averaged Navier–Stokes—on multiple generations of NVIDIA and AMD GPUs. It supports both structured and unstructured meshes and arbitrary polynomial orders, achieving superior throughput and strong scalability compared to state-of-the-art approaches.

Develop scalable GPU solvers for HDG discretizations of PDEsIntegrate preconditioners for efficient nonlinear PDE solutions on GPUsOptimize memory usage and parallelism with batched dense operations

Latest Papers

What's happening recently
View more

This work addresses the challenge of preconditioning indefinite or nonsymmetric sparse linear systems by proposing a novel multilevel preconditioner that integrates algebraic multigrid (AMG) hierarchies with graph neural networks (GNNs). For the first time, the method enables end-to-end learning of all-level smoothing, restriction, and interpolation operators within a unified framework, embedding AMG priors directly into the GNN architecture while remaining compatible with standard Krylov solvers. Extensive experiments on over 800 benchmark matrices demonstrate that the approach significantly accelerates convergence compared to classical AMG, ILUT, and existing GNN-based preconditioners on specific problem classes. The study also delineates the computational overhead introduced by the multilevel structure and extends the applicability of data-driven methods to non-elliptic problems.

algebraic multigridgraph neural networksiterative solvers

This work addresses the coarse-grid efficiency bottleneck in solving the pressure Poisson problem arising from large-scale incompressible Navier–Stokes simulations on exascale supercomputers. To overcome this challenge, the authors propose a novel two-level Schwarz method to replace conventional algebraic multigrid (AMG) solvers as the coarse-grid solver within a p-multigrid preconditioner. The method constructs a structured yet non-nested global coarse space that enables communication-free interpolation between the original p-multigrid coarse space and the new coarse representation, substantially enhancing parallel scalability. Implemented within the Nek5000/RS framework using spectral/finite element discretizations, the approach is evaluated on the Summit and Frontier supercomputers, demonstrating superior strong scaling and solution efficiency compared to state-of-the-art AMG solvers such as BoomerAMG.

coarse solverexascaleincompressible Navier-Stokes

This work addresses the longstanding challenge in algebraic multigrid (AMG) methods of balancing sparsity and convergence quality in coarse-grid operators. The authors propose RAPNet, the first graph neural network framework that directly learns nongalerkin coarse operators from sparse algebraic systems. Employing a hierarchical training strategy, RAPNet generalizes from small subgraphs to problems with millions of unknowns, achieving an effective trade-off among sparsity, convergence, and generalization while preserving computational efficiency during the solve phase. Experimental results demonstrate that RAPNet significantly outperforms classical nongalerkin baselines across a range of discretized partial differential equations and graph Laplacian problems, substantially accelerating multi-query tasks such as eigenvalue computations, time-dependent simulations, inverse problems, and design optimization.

algebraic multigridcoarse-grid operatorsscientific computing

Implicit time integration is key to robustly simulating stiff materials and large deformations, but its performance is often dominated by repeatedly solving large linear systems. Adaptive coarsening can reduce this cost by concentrating degrees of freedom (DoF) to where it is most needed, yet conventional explicit remeshing changes connectivity (and often vertex ordering), complicating parallel implementations, harming memory locality, and sometimes being disallowed when it may introduce local geometry intersections. Adaptive subspace approaches avoid topological changes, but basis construction and updates incur irregular data access patterns and typically produce dense system matrices, limiting GPU efficiency and keeping many practical systems CPU-centric. We present algebraic adaptive in-solve coarsening, a GPU-oriented method that dynamically reduces DoF within the Newton solve of implicit time integration without explicit topological modification. Starting from a fine mesh, we express adaptivity as a selective edge-collapse process governed by per-edge tags. Collapsible edges are aggregated in parallel using a warp-level hash mapping scheme that groups fine vertices into coarse super-nodes, while protected edges preserve local detail. This defines an implicit coarse mesh whose linear system is assembled algebraically by mapping and reducing fine-scale gradients and Hessians via efficient GPU reduction kernels. We solve the resulting coarse system with a preconditioned conjugate gradient (PCG) method and then prolongate the solution back to the fine mesh. Our approach integrates seamlessly with IPC's barrier energy and exploits GPU parallelism end-to-end. Across a range of challenging scenarios, we achieve up to 3x speedup over a state-of-the-art GPU IPC solver while producing visually indistinguishable results.

adaptive coarseningGPU accelerationimplicit time integration

This work addresses the limited efficiency of traditional algebraic multigrid (AMG) pressure solvers on unstructured grids, which stems from their sensitivity to grid irregularities. The study introduces, for the first time, a Graph Convolutional Isomorphism Network (GCIN) into the AMG framework to directly learn the algebraic structure of smoothers. By predicting optimal polynomial coefficients, the method constructs a sparse pseudoinverse operator that adaptively handles local anisotropy while preserving linear computational complexity. The proposed approach demonstrates strong generalization across scales and problem types, reducing V-cycle iteration counts and achieving practical speedups of 4%–37% on various benchmarks. Notably, it maintains robust convergence on grids up to 128 times larger than those used in training and on unseen industrial cases such as AirfRANS.

algebraic multigridcomputational bottleneckmesh irregularities

Hot Scholars

MK

Martin Kronbichler

Professor of Applied Numerics, Ruhr University Bochum
Finite element methodhigh performance computingcomputational fluid dynamicsmultigrid methods
BF

Bernat Font

TU Delft
CFDturbulencemachine learningHPC
UR

Ulrich Rüde

Professor for Computational Science and Engineering, FAU Erlangen-Nürnberg
Computational Science and EngineeringSupercomputingMultigridLattice Boltzmann Methods
MR

Matthias Rottmann

Professor of Computer Science, Osnabrück University, Germany
Computer VisionDeep LearningSafe AIEfficient AI
BZ

Bo Zhu

School of Interactive Computing, Georgia Institute of Technology
Computer GraphicsComputational PhysicsScientific Machine Learning