Score
Designs and builds experiments, benchmarks, load tests, and instrumentation to measure and validate system throughput, latency, scalability, and resource utilization under realistic workloads. Analyzes architectures, implementations, and configurations to identify bottlenecks, model capacity and performance trade-offs, and produce tuning, optimization, and capacity‑planning recommendations.
The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.
This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
This work addresses the challenge of efficiently conducting “What-If” I/O performance analysis for large-scale HPC applications, which is hindered by the complex interplay among access patterns, middleware, and file systems. The authors propose FBench, the first flexible I/O benchmarking tool based on context-free grammars (CFGs), capable of generating or replaying I/O traces—captured via Recorder—in real time without modifying application code. FBench supports both POSIX and MPI-IO interfaces and enables configuration-driven exploration through JSON-defined optimization strategies. It faithfully reproduces real-world workloads such as IOR, HACC-IO, FLASH Sedov, and LAMMPS. Evaluations on Lustre reveal that collective I/O write bandwidth can be up to 30× lower than ideal, burst buffers improve non-collective write bandwidth by 1.5×, and performance gains of up to 8× are achievable in LAMMPS scenarios, significantly accelerating I/O optimization studies.
Processor thermal design power (TDP) is widely misused as a proxy for actual power consumption in physics simulations, leading to inaccurate energy-efficiency assessments. Method: This study conducts the first empirical power and energy measurements of major production-scale physics simulation codes on heterogeneous exascale supercomputers at LLNL and Sandia. Leveraging multi-granularity energy modeling, cross-platform benchmarking, and real-time monitoring across commercial and advanced CPU–GPU heterogeneous nodes, it systematically quantifies runtime energy efficiency. Contribution/Results: Under typical simulation workloads, measured power draw is only 30–60% of TDP—substantially lower than nominal ratings. This work challenges the longstanding practice of substituting TDP for measured power, establishing an empirically grounded methodology for evaluating energy efficiency in exascale systems. It provides critical, reproducible, and generalizable energy benchmarks to guide hardware deployment and energy-aware optimization, thereby advancing low-carbon scientific computing.
This work addresses the high overhead of processing massive telemetry data in exascale supercomputing systems by proposing a heterogeneous acceleration–enabled, high-performance diagnostic framework. Integrating high-throughput C++ APIs with GPU-parallelized computation, the framework supports scalable MPI trace analysis and seamless integration with external tools. It introduces a novel topology-aware workflow that maps logical performance anomalies onto the physical coordinates of the Slingshot interconnect and pioneers a three-dimensional performance model to iteratively reconstruct application behavior, enabling precise identification of performance headroom. Evaluated on Aurora, the system ingests traces from 100,000 MPI ranks in just 9.69 seconds, achieving up to a 314× speedup over CPU-based analysis. On Frontier, it uncovers 32.28% potential acceleration for the GAMESS application.
Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.
This work addresses the absence of a systematic, traceable, and reproducible framework for reporting the performance of mathematical libraries—a gap that hinders accurate performance evaluation and resource planning for scientific applications on high-performance computing (HPC) systems. To this end, the paper introduces LAAB, the first framework explicitly designed around four core principles: traceability, compatibility, reliability, and accessibility. LAAB establishes an end-to-end reproducible performance evaluation pipeline through standardized benchmarking protocols, comprehensive metadata management, execution environment tracking, and advanced performance analysis techniques. The framework substantially enhances the accuracy and interoperability of mathematical library performance reporting, thereby providing a robust foundation for performance prediction and resource scheduling in scientific computing.