Performance Portable $\mathrm{SU}(N)$ Lattice Gauge Theory Simulation with Kokkos

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high maintenance costs of lattice gauge theory algorithms caused by the fragmentation of high-performance computing architectures. We present a performance-portable implementation of Wilson pure gauge Monte Carlo simulations built upon the Kokkos framework. By determining the SU(N) group order and spacetime dimensionality at compile time, the method maintains a single-source codebase that uniformly supports Serial, OpenMP, CUDA, HIP, and SYCL backends. Furthermore, communication overhead is optimized by overlapping MPI halo exchanges with interior updates. Experimental results demonstrate that this implementation accurately reproduces physical observables while achieving performance comparable to native code on NVIDIA A100 GPUs and delivering SIMD acceleration on Armv9 processors. Large-scale benchmarks on the LineShine supercomputer confirm significant speedups, validating the effectiveness of our approach.
📝 Abstract
The increasing diversity of high performance computing systems makes separate, architecture specific implementations of lattice gauge theory algorithms costly to maintain. We present \texttt{kwqft}, a performance portable Kokkos implementation of Wilson pure gauge Monte Carlo simulation for $\mathrm{SU}(N)$ Yang-Mills theory in an arbitrary number of space-time dimensions. The gauge group order $N$ and the dimension $D$ are compile time parameters. A single source targets the Serial, OpenMP, CUDA, HIP, and SYCL execution spaces, with MPI halo exchange overlapped with interior updates. The implementation reproduces the exact two-dimensional plaquette and published three and four dimensional values for gauge groups up to $\mathrm{SU}(17)$. On an NVIDIA A100 the Kokkos CUDA backend is competitive with a native CUDA code, SIMD acceleration improves the OpenMP path on Armv9 processors, and a large scale speedup is demonstrated for an $\mathrm{SU}(4)$ lattice on the LineShine supercomputer, currently ranked first on the TOP500 list.
Problem

Research questions and friction points this paper is trying to address.

Lattice Gauge Theory
Performance Portability
High Performance Computing
SU(N) Yang-Mills
Innovation

Methods, ideas, or system contributions that make the work stand out.

Performance Portability
Lattice Gauge Theory
Kokkos
SU(N) Yang-Mills
Monte Carlo Simulation
🔎 Similar Papers
No similar papers found.