Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2

📅 2026-06-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of performance and energy efficiency for cutting-edge scientific applications on emerging heterogeneous supercomputing nodes, particularly those featuring CPU+GPU协同 architectures. For the first time, we conduct fine-grained benchmarking of five representative scientific workloads—spanning molecular dynamics, astrophysics, and finite-element PDE solvers—on SuperMUC-NG Phase 2 nodes equipped with Intel Ponte Vecchio GPUs, leveraging the lightweight power monitoring tool p3em and the Energy Aware Runtime (EAR). Our results demonstrate that GPU acceleration yields throughput improvements of 4–12× and up to 15× higher energy efficiency, most notably for LAMMPS and AthenaK, though these gains diminish with smaller problem sizes. Additionally, we observe that CPU-only executions consistently underutilize the node’s thermal design power, revealing significant headroom for runtime and scheduling optimizations.
📝 Abstract
We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (LRZ) equipped with Intel Xeon Platinum 8480+ and Intel Data Center GPU Max 1550 (Ponte Vecchio, PVC) accelerators. The selected codes span molecular dynamics (gromacs, lammps), astrophysics and cosmology (OpenGadget3, AthenaK), and finite-element PDE solvers from the dealii-X Center of Excellence. For each code we measure throughput and energy efficiency expressed as compute-elements per wall-clock second (or per Joule of consumed energy) on a single compute node, comparing CPU-only (SPR) against combined CPU+GPU (SPR+PVC) configurations where available. Energy measurements rely on lightweight code instrumentation with p3em, or the Energy Aware Runtime (EAR) present on the system. Our results show that GPU offload yields $4-12\times$ higher throughput and up to $15\times$ better energy efficiency compared to CPU-only execution, with lammps and AthenaK benefiting most. However, both throughput and energy gains are sensitive to problem granularity: insufficient work per GPU tile erodes the accelerator advantage, as clearly observed in AthenaK at small mesh-block sizes. The power-budget utilization is systematically lower for CPUs than it is for GPUs, indicating that even at peak useful-work rate, most applications running on CPUs leave a significant fraction of the node's thermal envelope unused.
Problem

Research questions and friction points this paper is trying to address.

performance characterization
energy efficiency
scientific applications
CPU-GPU systems
supercomputing
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU acceleration
energy efficiency
performance characterization
HPC workloads
power budget utilization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Salvatore Cielo
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
E
Elmira Birang
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
A
Alexander Pöppl
Intel Deutschland GmbH, Germany
S
Sajad Azizi
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
P
Plamen Dobrev
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
M
Margarita Egelhofer
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
I
Ivan Pribec
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany
G
Gerald Mathias
Leibniz Supercomputing Center (LRZ), Boltzmannstraße 1, 85748 Garching b.München, Germany