Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work challenges the conventional view of MPI communication blocking as a performance bottleneck by exploring its potential to mitigate memory bandwidth contention. Through an architecture-aware desynchronization strategy, compute- and memory-intensive kernels are overlapped, transforming communication stalls into opportunities for bandwidth optimization. The study further reveals that indiscriminate asynchronous communication can exacerbate contention and prove counterproductive. Employing micro-benchmarks, noise injection, and a bandwidth-aware simulator, the proposed approach achieves significant speedups in an optical flow solver. Additionally, it identifies the optimal number of concurrent processes required to balance bandwidth saturation points within ccNUMA domains.
📝 Abstract
Parallel applications are often designed for synchronous, lock-step execution, treating communication stalls as performance hazards. Yet, in a communication-light application without frequent synchronization points that alternates between compute-bound memory-bound execution, an MPI communication stall can act as an unintentional relief on memory-bandwidth contention. We demonstrate this using a Parallel Optical Flow Solver, which combines a compute-bound Ray Tracing kernel with a memory-bound Optical Flow Solver kernel and negligible inter-process communication. This program shows considerable speedup via desynchronization and automatic overlap between compute- and memory-bound phases, showing that natural desynchronization is an architecture-aware optimization. An optimal speedup is achieved when the number of processes concurrently executing the memory-bound phase on a ccNUMA domain is near the bandwidth saturation point. We also show a case where reducing communication overhead using MPI asynchronous progress significantly degrades performance because it allows too many ranks to contend for memory bandwidth simultaneously. In order to study the dynamics under more controlled conditions, we develop a tunable dual-kernel microbenchmark, with which we show that significant application or system noise (natural or injected) is required to achieve full desynchronization. Finally, we also validate these results using a bandwidth-aware, model-based simulator.
Problem

Research questions and friction points this paper is trying to address.

MPI applications
compute-bound kernels
memory-bound kernels
desynchronization
memory bandwidth contention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Desynchronization
Memory-bandwidth contention
Compute- and memory-bound kernels
MPI communication stall
Microbenchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Ayesha Afzal
Erlangen National High Performance Computing Center (NHR@FAU), Erlangen, Germany
K
Krishna Manda
University of Bonn, Germany
Georg Hager
Georg Hager
Friedrich-Alexander-Universität Erlangen-Nürnberg
High Performance Computing