device-resident blocked amg

Designs and implements algebraic multigrid solvers for block-structured systems that execute entirely on accelerator devices (device-resident), ensuring all setup and solve phases remain on the device. Work includes smoothed-aggregation or other AMG setup in block form, constructing blocked prolongators, computing Galerkin products, assembling blocked coarse operators (e.g., in COO), and implementing device-resident V-cycle and multigrid operations.

device-residentblockedamg

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiency of existing GPU-based algebraic multigrid (AMG) implementations that rely on scalar sparse matrix formats, which fail to exploit the natural rectangular block structure arising from discretizations of vector-valued PDEs such as elasticity, leading to high memory overhead and low arithmetic intensity. The paper presents the first fully GPU-resident, native block AMG pathway in PETSc, built upon Kokkos to support a portable block-sparse matrix type with heterogeneous block sizes across coarse and fine grids while preserving block structure throughout the smoothed aggregation AMG process. It integrates PetscSF for block prolongation communication across processes and introduces a dense rectangular block COO assembly mechanism. Experiments on 3D elasticity problems using A100 GPUs demonstrate significantly reduced GPU memory usage compared to cuSPARSE and scalar Kokkos approaches, achieving a 2.27× speedup in Galerkin construction at 64-GPU scale, along with 1.24× and 1.42× improvements in V-cycle and SpMV performance, respectively.

algebraic multigridblocked matriceselasticity

This work addresses the high data movement overhead and lack of sparse library support in block algebraic multigrid (AMG) due to the nonstandard block structure of the Galerkin triple product $A_c = P^T A P$. The authors propose shared-memory block kernels and a search-free scheduling strategy, along with a novel prolongation operator filtering algorithm based on the Frobenius norm, which significantly reduces data transfer, coarse-operator fill-in, and memory footprint while preserving iterative convergence. A fully GPU-resident block AMG pipeline is implemented using Kokkos (with CUDA backend) and native CUDA, eliminating scalar expansion and host-device data copies. On an NVIDIA A100, DRAM traffic for the fine-grid Galerkin product drops from 17.4 GB to 10.5 GB, execution time decreases from 82 ms to 45 ms, and hotspot operations achieve a 2.9× speedup, approaching theoretical bandwidth limits.

algebraic multigridblock sparse matricesdata movement

Extending DD-$alpha$AMG on heterogeneous machines

Jul 10, 2024
LH
Lianhua He
🏛️ Computer Network Information Center, Chinese Academy of Sciences | Department of Mathematics, Bergische Universität Wuppertal

This work addresses the key challenge of enabling efficient, cross-architecture execution of the DD-αAMG multigrid solver for lattice QCD simulations on heterogeneous exascale supercomputers (e.g., ORISE). Method: We present the first full port of DD-αAMG to the HIP platform, unifying acceleration across both NVIDIA and AMD GPUs; introduce an improved Richardson smoother that achieves superior convergence–speed trade-offs—outperforming GCR significantly and lagging SAP by only ~10%; and extend the odd-even preconditioned GMRES/Richardson hybrid iteration, applying advanced SIMD vectorization and computational restructuring to coarse-grid operations. Contribution/Results: Experiments on ORISE demonstrate efficient heterogeneous acceleration, with substantial performance gains in critical coarse-grained operations. The proposed framework establishes a scalable, cross-platform AMG solver paradigm tailored for large-scale lattice QCD computations.

Extend DD-αAMG for heterogeneous machinesImprove smoothers like Richardson in DD-αAMGPort coarse-grid operations via advanced vectorization

This work addresses the coarse-grid efficiency bottleneck in solving the pressure Poisson problem arising from large-scale incompressible Navier–Stokes simulations on exascale supercomputers. To overcome this challenge, the authors propose a novel two-level Schwarz method to replace conventional algebraic multigrid (AMG) solvers as the coarse-grid solver within a p-multigrid preconditioner. The method constructs a structured yet non-nested global coarse space that enables communication-free interpolation between the original p-multigrid coarse space and the new coarse representation, substantially enhancing parallel scalability. Implemented within the Nek5000/RS framework using spectral/finite element discretizations, the approach is evaluated on the Summit and Frontier supercomputers, demonstrating superior strong scaling and solution efficiency compared to state-of-the-art AMG solvers such as BoomerAMG.

coarse solverexascaleincompressible Navier-Stokes

AMReX: Block-structured adaptive mesh refinement for multiphysics applications

Sep 25, 2020
WZ
Weiqun Zhang
🏛️ Lawrence Berkeley National Laboratory

Efficiently implementing block-structured adaptive mesh refinement (AMR) for multi-physics simulations—such as accelerator design, combustion, and cosmology—across heterogeneous hardware platforms (from laptops to exascale supercomputers) faces persistent challenges in numerical accuracy, scalability, and hardware adaptability. This paper introduces the first general-purpose AMR infrastructure unifying data containers, parallel iterators, and hardware-aware schedulers, enabling a unified software framework that supports seamless co-execution and heterogeneous deployment (CUDA, HIP, Kokkos, MPI, OpenMP) of PDE-based, particle-based, and particle-mesh coupled algorithms. Built upon a nested-grid abstraction, the framework decouples geometry, algorithmic logic, and backend execution. It achieves >90% strong and weak scaling efficiency on million-core systems, significantly reducing memory footprint and computational overhead. The infrastructure has already enabled over ten applications within the U.S. Exascale Computing Project (ECP).

Develops AMReX for efficient multiphysics simulations.Reduces computational cost with adaptive mesh refinement.Supports diverse applications from laptops to exascale.

Latest Papers

What's happening recently
View more

This work addresses the limited efficiency of traditional algebraic multigrid (AMG) pressure solvers on unstructured grids, which stems from their sensitivity to grid irregularities. The study introduces, for the first time, a Graph Convolutional Isomorphism Network (GCIN) into the AMG framework to directly learn the algebraic structure of smoothers. By predicting optimal polynomial coefficients, the method constructs a sparse pseudoinverse operator that adaptively handles local anisotropy while preserving linear computational complexity. The proposed approach demonstrates strong generalization across scales and problem types, reducing V-cycle iteration counts and achieving practical speedups of 4%–37% on various benchmarks. Notably, it maintains robust convergence on grids up to 128 times larger than those used in training and on unseen industrial cases such as AirfRANS.

algebraic multigridcomputational bottleneckmesh irregularities

Implicit time integration is key to robustly simulating stiff materials and large deformations, but its performance is often dominated by repeatedly solving large linear systems. Adaptive coarsening can reduce this cost by concentrating degrees of freedom (DoF) to where it is most needed, yet conventional explicit remeshing changes connectivity (and often vertex ordering), complicating parallel implementations, harming memory locality, and sometimes being disallowed when it may introduce local geometry intersections. Adaptive subspace approaches avoid topological changes, but basis construction and updates incur irregular data access patterns and typically produce dense system matrices, limiting GPU efficiency and keeping many practical systems CPU-centric. We present algebraic adaptive in-solve coarsening, a GPU-oriented method that dynamically reduces DoF within the Newton solve of implicit time integration without explicit topological modification. Starting from a fine mesh, we express adaptivity as a selective edge-collapse process governed by per-edge tags. Collapsible edges are aggregated in parallel using a warp-level hash mapping scheme that groups fine vertices into coarse super-nodes, while protected edges preserve local detail. This defines an implicit coarse mesh whose linear system is assembled algebraically by mapping and reducing fine-scale gradients and Hessians via efficient GPU reduction kernels. We solve the resulting coarse system with a preconditioned conjugate gradient (PCG) method and then prolongate the solution back to the fine mesh. Our approach integrates seamlessly with IPC's barrier energy and exploits GPU parallelism end-to-end. Across a range of challenging scenarios, we achieve up to 3x speedup over a state-of-the-art GPU IPC solver while producing visually indistinguishable results.

adaptive coarseningGPU accelerationimplicit time integration

This work addresses the significant performance degradation of high-order finite element stiffness operators in end-to-end scenarios due to inefficient utilization of modern processor matrix engines. Focusing on the stiffness operator in SPECFEM3D running on Arm LX2 CPUs, the study proposes a holistic operator-level co-optimization methodology that transcends conventional approaches limited to optimizing tensor contraction kernels alone. By systematically co-designing pointwise computations, field data layouts, and coefficient memory access patterns—through explicit SIMD optimization, data relayout, and vectorized blocked streaming loads—the approach achieves a matrix engine speedup of 1.6× for the full stiffness operator, up from 1.1×, approaching the theoretical upper bound. Factorization-based diagnostics and contraction-free ablation experiments validate the efficacy and necessity of this full-path co-optimization strategy.

finite elementsmatrix enginesperformance gap

This work addresses the scalability bottleneck of large-scale incompressible fluid simulations on multi-node CPU/GPU clusters by proposing an MPI-based heterogeneous parallel optimization framework integrated with an enhanced geometric multigrid Poisson solver. The core innovations include an adaptive under-relaxed red-black Gauss–Seidel smoother and an anisotropic coarsening operator, both implemented purely in Julia. Implemented within WaterLily.jl, the method achieves near-ideal strong scaling and weak scaling efficiency exceeding 85%. Notably, it attains over 96% weak scaling efficiency across nodes at billion-cell resolution, substantially enhancing solver performance and memory concurrency.

geometric multigridincompressible flowMPI parallelism

This work addresses the challenges of complex boundary condition handling, code redundancy, and poor distributed efficiency in partial differential equation solvers on block-structured grids. The authors propose a unified modeling approach that expresses user-defined boundary conditions as affine sparse linear operators and, for the first time, systematically reformulates them into sparse matrix-vector multiplication (SpMV) form. Leveraging a domain-specific language (DSL) and compiler techniques—combined with multi-stage programming and polyhedral analysis—the framework automatically generates highly optimized matrix-free or sparse matrix kernels while optimizing communication scheduling and reuse. The method achieves substantial performance gains, demonstrating 72%–88% strong scaling efficiency on 1,344 CPU cores, up to 7.6× acceleration in boundary computation kernels, and a reduction of over 70% in code size.

block-structured gridsboundary conditiondistributed-memory execution