Score
Designs, implements, and evaluates systems, algorithms, and runtime strategies that split and coordinate computational workloads across processors, accelerators, and nodes — including data, model, pipeline, sequence, expert, and multi‑dimensional (e.g., 5D) parallelism — using parallel programming models and distributed/parallel computing frameworks. Builds and tunes parallel algorithms and communication/topology schemes, and analyzes performance, scalability, resource utilization, and synchronization/communication overhead to select and optimize parallelization strategies and techniques.
Selecting appropriate parallel programming models for heterogeneous HPC architectures remains challenging due to divergent hardware characteristics and software trade-offs. Method: This paper conducts the first multi-dimensional quantitative comparison of MPI, OpenMP, and CUDA—evaluating architectural adaptability, scalability bottlenecks, development complexity, and domain suitability—and proposes a hybrid programming model selection framework tailored to heterogeneity. The framework integrates communication modeling, memory contention analysis, and GPU kernel optimization for empirical validation. Contribution/Results: Experiments show MPI achieves >92% strong scaling efficiency in distributed, communication-intensive workloads; OpenMP delivers 3.8× speedup on shared-memory loop-parallel tasks; CUDA attains up to 12.5× acceleration on data-parallel kernels; and hybrid strategies yield an average 27% improvement in end-to-end performance. The study provides both theoretical foundations and practical guidelines for optimizing and co-designing programming models in heterogeneous HPC environments.
To address the high overhead of dynamic data repartitioning in multi-core HPC systems under time-varying workloads, this paper proposes a lightweight, hierarchical partitioning method jointly driven by geometric and statistical principles. The method integrates space-filling curve ordering, greedy knapsack-based load balancing, and hierarchical data decomposition to support efficient dynamic partitioning of 2D/3D structured grids, point sets, and general graphs. It introduces, for the first time, an adaptive repartitioning mechanism guided by real-time feedback on data distribution, substantially reducing computational and communication overhead in frequently updated scenarios. Implemented via a hybrid parallel programming model (MPI + OpenMP) on modern many-core architectures, experimental results demonstrate a 3.2–5.7× speedup in partitioning time and a load imbalance ratio below 3.1%. This approach provides timely, low-overhead data partitioning support for parallel algorithms in large-scale scientific computing.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
To address the challenges of cross-platform orchestration and fragmented resource scheduling in hybrid HPC–ML workflows, this paper proposes a service-oriented, scalable runtime architecture. Building upon the RADICAL-Pilot framework, we introduce the first service-oriented execution model enabling dynamic, multi-granularity, low-overhead coordination of heterogeneous HPC and ML tasks. Our approach unifies resource abstraction across platforms, implements distributed task scheduling, and jointly orchestrates AI and HPC workloads—thereby enabling seamless coupling and coordinated scheduling between on-premises exascale supercomputers and cloud environments. Experimental evaluation on an exascale prototype system demonstrates concurrent deployment of multiple ML models with runtime overhead under 2%. The architecture successfully supports three representative data-driven scientific applications, effectively overcoming the traditional siloing of HPC and ML workflows.
Existing parallel computing curricula for undergraduate and graduate students often lack a unified, principle-centered pedagogical framework that balances theoretical foundations with practical implementation while ensuring broad applicability. Method: This work develops a systematic lecture note suite grounded in deterministic parallel algorithms, covering core theory (work-time model, efficiency and scalability analysis), mainstream programming models (OpenMP, MPI, pthreads), and C-language implementation—explicitly excluding GPU programming and randomized algorithms to preserve conceptual generality. It integrates visualization-guided explanations, verifiable code examples, and structured programming exercises emphasizing universal performance criteria: execution time, energy consumption, and scalability. Contribution/Results: The resulting self-contained, production-ready lecture notes are accompanied by open-source code and extensible problem sets. They effectively support both formal instruction in parallel and high-performance computing courses and independent learning, enhancing pedagogical coherence and practical accessibility.
This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.
This study systematically investigates hybrid parallelism strategies for large language models during both training and inference, aiming to balance computational, communication, and memory overheads. By constructing a mathematical cost model grounded in collective communication operations and integrating communication-computation overlap with automated strategy search, the work proposes a hybrid parallelism framework that achieves both efficiency and scalability. It is the first to unify theoretical modeling, automated search, and empirical evaluation across multiple hardware architectures, revealing the trade-offs among different parallelization strategies in training versus inference. The resulting framework provides reusable deployment guidelines for canonical model architectures, significantly enhancing distributed efficiency.
To address communication bottlenecks—both on-chip and inter-node—in post-exascale supercomputers and AI data centers, this paper proposes a unified, scalable interconnect architecture spanning chip-level and system-level hierarchies. Methodologically, it introduces a novel low-diameter network topology, fine-grained flow control mechanisms, and heterogeneous resource co-sharing strategies, tightly integrated with high-bandwidth memory and accelerator hardware. Its key contribution lies in jointly optimizing communication latency and bandwidth, substantially alleviating resource contention and improving data locality. Experimental evaluation at scale—up to 1,000 accelerators—demonstrates a 32–47% reduction in communication overhead, a 2.1× increase in system throughput, and a 38% improvement in energy efficiency. The architecture thus delivers efficient, scalable interconnect support for generative AI workloads and large-scale scientific simulations.
This study addresses the challenges AI-driven scientific workflows encounter in high-performance computing (HPC) environments—namely data intensity, resource heterogeneity, frequent iteration cycles, and I/O bottlenecks—which hinder their compatibility with conventional linear pipeline architectures. The work presents the first systematic integration of AI workflow characteristics with HPC system design, proposing a transformative framework tailored for adaptive intelligent computing environments and articulating twelve practical design principles. This framework incorporates key technologies including containerization, job array scheduling, explicit feedback mechanisms, heterogeneous resource management, and small-file I/O optimization, making it particularly well-suited for high-throughput domains such as computational biology. The resulting guidelines offer researchers actionable strategies to substantially enhance the efficiency, portability, and scalability of AI-HPC workflows.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.