high-performance computing

Designs, builds, and analyzes parallel and distributed computing systems and software to achieve high throughput and low-latency execution on multicore CPUs, GPUs, accelerators, clusters, and supercomputers. Work includes writing and optimizing scalable code, tuning memory and I/O, profiling and benchmarking performance, and configuring schedulers, interconnects, and runtime environments to meet scalability and efficiency requirements.

high-performancecomputing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An HPC Benchmark Survey and Taxonomy for Characterization

Sep 10, 2025
AH
Andreas Herten
🏛️ Forschungszentrum Jülich | Lawrence Livermore National Laboratory | Texas A&M University

The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.

Developing taxonomy to categorize HPC benchmarks systematicallyProviding structured comparison of hardware and software evaluation toolsSurveying existing HPC benchmarks for comprehensive characterization

Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey

Dec 31, 2024
HA
Hammad Ather
🏛️ University of Oregon | Lawrence Berkeley National Laboratory | Lawrence Livermore National Laboratory | The Ohio State University

With AI and high-resolution simulations increasingly driving HPC workloads, parallel I/O performance bottlenecks have grown more complex, while existing optimization tools remain fragmented and difficult to select. Method: We systematically review 131 publications and—employing bibliometric analysis, systematic literature review, and taxonomy modeling—construct the first comprehensive, end-to-end parallel I/O classification framework (a “360° taxonomy”) covering characterization, analysis, and optimization. Our approach integrates cross-platform profiling and tracing tools—including Darshan, Vampir, and Lustre trace—into a unified analytical pipeline. Contribution: We propose the first holistic, cross-layer I/O optimization framework spanning applications, runtime systems, file systems, and hardware; release a structured knowledge graph and open-source classification toolkit; and significantly reduce decision-making overhead in selecting optimization strategies. This work delivers a reusable, scalable methodology for enhancing parallel I/O performance in production HPC environments.

I/O OptimizationParallel I/OSupercomputer Systems

Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems

Mar 13, 2025
FK
Fabian Knorr
🏛️ University of Innsbruck

SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.

Enhancing memory allocation and concurrency in SYCL programs on accelerator clusters.Optimizing scheduling for high-level parallel programs on multi-GPU systems.Reducing delays in distributed-memory applications through graph-based representations.

Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA

Jun 18, 2025
NA
Nizar ALHafez
🏛️ Higher Institute for Applied Sciences and Technology | Damascus University

Selecting appropriate parallel programming models for heterogeneous HPC architectures remains challenging due to divergent hardware characteristics and software trade-offs. Method: This paper conducts the first multi-dimensional quantitative comparison of MPI, OpenMP, and CUDA—evaluating architectural adaptability, scalability bottlenecks, development complexity, and domain suitability—and proposes a hybrid programming model selection framework tailored to heterogeneity. The framework integrates communication modeling, memory contention analysis, and GPU kernel optimization for empirical validation. Contribution/Results: Experiments show MPI achieves >92% strong scaling efficiency in distributed, communication-intensive workloads; OpenMP delivers 3.8× speedup on shared-memory loop-parallel tasks; CUDA attains up to 12.5× acceleration on data-parallel kernels; and hybrid strategies yield an average 27% improvement in end-to-end performance. The study provides both theoretical foundations and practical guidelines for optimizing and co-designing programming models in heterogeneous HPC environments.

Analyze performance, suitability, complexity of parallel programming approachesCompare MPI, OpenMP, CUDA for HPC parallel programming modelsEvaluate hybrid models for optimal HPC application performance

Lectures on Parallel Computing

Jul 26, 2024
JT
J. Träff
🏛️ TU Wien

Existing parallel computing curricula for undergraduate and graduate students often lack a unified, principle-centered pedagogical framework that balances theoretical foundations with practical implementation while ensuring broad applicability. Method: This work develops a systematic lecture note suite grounded in deterministic parallel algorithms, covering core theory (work-time model, efficiency and scalability analysis), mainstream programming models (OpenMP, MPI, pthreads), and C-language implementation—explicitly excluding GPU programming and randomized algorithms to preserve conceptual generality. It integrates visualization-guided explanations, verifiable code examples, and structured programming exercises emphasizing universal performance criteria: execution time, energy consumption, and scalability. Contribution/Results: The resulting self-contained, production-ready lecture notes are accompanied by open-source code and extensible problem sets. They effectively support both formal instruction in parallel and high-performance computing courses and independent learning, enhancing pedagogical coherence and practical accessibility.

Covering OpenMP and MPI frameworks for parallel programmingFocusing on deterministic algorithms for shared/distributed memory systemsIntroducing theoretical concepts for analyzing parallel algorithms

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic evaluation of performance and energy efficiency for cutting-edge scientific applications on emerging heterogeneous supercomputing nodes, particularly those featuring CPU+GPU协同 architectures. For the first time, we conduct fine-grained benchmarking of five representative scientific workloads—spanning molecular dynamics, astrophysics, and finite-element PDE solvers—on SuperMUC-NG Phase 2 nodes equipped with Intel Ponte Vecchio GPUs, leveraging the lightweight power monitoring tool p3em and the Energy Aware Runtime (EAR). Our results demonstrate that GPU acceleration yields throughput improvements of 4–12× and up to 15× higher energy efficiency, most notably for LAMMPS and AthenaK, though these gains diminish with smaller problem sizes. Additionally, we observe that CPU-only executions consistently underutilize the node’s thermal design power, revealing significant headroom for runtime and scheduling optimizations.

CPU-GPU systemsenergy efficiencyperformance characterization

This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.

code representationcomputational kernelsHPC benchmarks

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the high complexity of existing high-performance computing (HPC) performance analysis tools, which hinders students’ intuitive understanding of parallel program performance issues. To bridge this gap, the paper introduces EduMPI—the first educational tool that integrates HPC cluster operations and MPI performance analysis within a streamlined graphical interface. EduMPI enables near real-time, physically node-layout-aware communication visualization, facilitating interactive identification of load imbalance and other performance bottlenecks. User studies demonstrate that, compared to professional-grade tools, EduMPI significantly lowers the learning barrier and effectively enhances students’ comprehension of parallel performance characteristics, thereby improving the practicality and accessibility of parallel programming education.

educational integrationMPIparallel programming education

Hot Scholars

FA

Faez Ahmed

Associate Professor, MIT
Generative AIEngineering DesignMachine LearningEngineering Optimization
SS

Stephan Simonis

Karlsruhe Institute of Technology (KIT)
exploratory computationcomputational mathematicsscientific computingapplied mathematics
ME

Mohamed Elrefaie

Massachusetts Institute of Technology
AerodynamicsCFDMachine LearningGenerative AI
MZ

Mingliang Zhong

Karlsruhe Institute of Technology
computational fluid dynamics
JY

Jong Youl Choi

Oak Ridge National Laboratory
big data sciencedata intensive computingdata mining