Score
Design, implement, and analyze parallel and distributed software that combines message-passing (MPI), shared-memory multithreading (OpenMP), and accelerator/GPU offloading to decompose computation across processes, threads, and device warps, manage inter-node and intra-node communication, and map computation to physical nodes and accelerators. Optimize runtime and memory usage, ensure correct concurrency and synchronization (including warp-synchronous and thread-lane patterns), and evaluate scaling, load balance, and performance bottlenecks using profiling and communication-visualization techniques.
Selecting appropriate parallel programming models for heterogeneous HPC architectures remains challenging due to divergent hardware characteristics and software trade-offs. Method: This paper conducts the first multi-dimensional quantitative comparison of MPI, OpenMP, and CUDA—evaluating architectural adaptability, scalability bottlenecks, development complexity, and domain suitability—and proposes a hybrid programming model selection framework tailored to heterogeneity. The framework integrates communication modeling, memory contention analysis, and GPU kernel optimization for empirical validation. Contribution/Results: Experiments show MPI achieves >92% strong scaling efficiency in distributed, communication-intensive workloads; OpenMP delivers 3.8× speedup on shared-memory loop-parallel tasks; CUDA attains up to 12.5× acceleration on data-parallel kernels; and hybrid strategies yield an average 27% improvement in end-to-end performance. The study provides both theoretical foundations and practical guidelines for optimizing and co-designing programming models in heterogeneous HPC environments.
To address portability bottlenecks in HPC arising from GPU heterogeneity and challenging distributed memory management, this paper introduces DiOMP, a distributed OpenMP framework. Methodologically, DiOMP features a novel unified runtime that integrates OpenMP target offloading with Partitioned Global Address Space (PGAS) semantics, enabling both symmetric and asymmetric GPU memory allocation. It further incorporates OMPCCL—a lightweight, portable collective communication layer—and is implemented via LLVM/OpenMP extensions, supporting GASNet-EX and GPI-2 communication backends across NVIDIA, AMD, and Grace Hopper platforms. Evaluation on A100, Grace Hopper, and MI250X systems demonstrates significant performance improvements for applications including matrix multiplication and MiniMod, while achieving strong scalability and programming simplicity.
This study addresses the challenge of achieving both cross-architecture portability and high performance in single-node, multi-GPU scientific computing across NVIDIA, AMD, and Intel platforms. Focusing on a 3D heat conduction problem, it presents the first unified implementation of multi-GPU collaboration using OpenMP offloading across all three major GPU architectures, and systematically evaluates its communication efficiency, memory management, and scalability relative to native programming models—namely CUDA, HIP, and SYCL. Experimental results demonstrate that OpenMP offloading achieves approximately 2× and 4× speedup on dual-GPU and quad-GPU configurations, respectively, confirming its potential to deliver near-native performance while maintaining strong portability across heterogeneous hardware ecosystems.
This work addresses the high overhead of processing massive telemetry data in exascale supercomputing systems by proposing a heterogeneous acceleration–enabled, high-performance diagnostic framework. Integrating high-throughput C++ APIs with GPU-parallelized computation, the framework supports scalable MPI trace analysis and seamless integration with external tools. It introduces a novel topology-aware workflow that maps logical performance anomalies onto the physical coordinates of the Slingshot interconnect and pioneers a three-dimensional performance model to iteratively reconstruct application behavior, enabling precise identification of performance headroom. Evaluated on Aurora, the system ingests traces from 100,000 MPI ranks in just 9.69 seconds, achieving up to a 314× speedup over CPU-based analysis. On Frontier, it uncovers 32.28% potential acceleration for the GAMESS application.
Existing parallel computing curricula for undergraduate and graduate students often lack a unified, principle-centered pedagogical framework that balances theoretical foundations with practical implementation while ensuring broad applicability. Method: This work develops a systematic lecture note suite grounded in deterministic parallel algorithms, covering core theory (work-time model, efficiency and scalability analysis), mainstream programming models (OpenMP, MPI, pthreads), and C-language implementation—explicitly excluding GPU programming and randomized algorithms to preserve conceptual generality. It integrates visualization-guided explanations, verifiable code examples, and structured programming exercises emphasizing universal performance criteria: execution time, energy consumption, and scalability. Contribution/Results: The resulting self-contained, production-ready lecture notes are accompanied by open-source code and extensible problem sets. They effectively support both formal instruction in parallel and high-performance computing courses and independent learning, enhancing pedagogical coherence and practical accessibility.
This work addresses the performance limitations of current GPU communication APIs, which either rely on CPU involvement or impose substantial synchronization overhead, thereby constraining the efficiency of machine learning and high-performance computing applications. The authors propose and implement a novel MPI-based GPU communication abstraction that, for the first time, enables fully CPU-bypassed GPU-to-GPU communication within the MPI framework and natively supports halo exchange primitives such as gather and scatter. By integrating MPI extensions, HPE Slingshot 11 network hardware, and the Cabana/Kokkos portable programming model, the design achieves a 50% reduction in medium-message latency and demonstrates a 28% improvement in halo exchange performance at strong scale on 8,192 GPUs on the Frontier supercomputer.
This study systematically examines the four-decade evolution of synchronization mechanisms and in-network computing architectures in large-scale parallel systems. Tracing the trajectory from the NYU Ultracomputer to modern exascale supercomputers, it integrates key technological milestones—including Fetch-and-Add, multistage interconnection networks, MPI, PCIe atomics, GPU cache coherence mappings, and HIP/Triton compilation stacks—to uncover, for the first time, the dynamic interplay and competition among shared-memory, message-passing, and in-network computing paradigms. The work elucidates the continuous co-evolution of synchronization primitives across hardware-software boundaries, offering critical historical context and architectural insights for the design of future heterogeneous supercomputing systems.
This study addresses the performance bottleneck in tiled Cholesky decomposition under non-uniform workloads caused by implicit synchronization barriers inherent in the fork-join execution model. The authors introduce Cholesky-Bench, a benchmark to systematically evaluate four parallelization strategies—classical fork-join, loop-fused fork-join, synchronous tasks, and asynchronous tasks based on explicit data dependencies—implemented using both OpenMP and HPX on a dual-socket 128-core AMD Zen 2 system. For the first time, they quantitatively compare the practical benefits of Asynchronous Many-Task (AMT) runtimes against optimized fork-join models on irregular kernels, reveal the impact of compiler optimizations on task overhead, and propose an effective optimization that eliminates redundant synchronization. Experimental results show that HPX achieves 15%–30% higher performance than OpenMP at optimal tile sizes, with asynchronous HPX tasks outperforming their OpenMP counterparts by 26%, reducing task overhead by 3.8×, and yielding an additional 7%–14% speedup through redundant synchronization removal.
This work proposes OMPDataPerf, a novel framework that addresses performance bottlenecks in heterogeneous OpenMP applications caused by inefficient data mappings—a problem that existing tools struggle to diagnose automatically. Leveraging the OpenMP Tools Interface (OMPT), OMPDataPerf enables the first automated dynamic detection and profiling of data transfer and allocation patterns in heterogeneous OpenMP programs. By integrating runtime instrumentation, dynamic analysis, and performance modeling, the approach precisely identifies source code locations responsible for suboptimal data mappings, estimates their optimization potential, and delivers actionable recommendations—all with only a 5% geometric mean runtime overhead. This significantly reduces the manual effort traditionally required for performance diagnosis and optimization in heterogeneous parallel programming.
This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.