performance engineering

Designs and builds experiments, benchmarks, load tests, and instrumentation to measure and validate system throughput, latency, scalability, and resource utilization under realistic workloads. Analyzes architectures, implementations, and configurations to identify bottlenecks, model capacity and performance trade-offs, and produce tuning, optimization, and capacity‑planning recommendations.

performanceengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$231K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An HPC Benchmark Survey and Taxonomy for Characterization

Sep 10, 2025
AH
Andreas Herten
🏛️ Forschungszentrum Jülich | Lawrence Livermore National Laboratory | Texas A&M University

The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.

Developing taxonomy to categorize HPC benchmarks systematicallyProviding structured comparison of hardware and software evaluation toolsSurveying existing HPC benchmarks for comprehensive characterization

This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.

cluster-wide benchmarkingcontinuous testingdata center validation

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

This work addresses the challenge of efficiently conducting “What-If” I/O performance analysis for large-scale HPC applications, which is hindered by the complex interplay among access patterns, middleware, and file systems. The authors propose FBench, the first flexible I/O benchmarking tool based on context-free grammars (CFGs), capable of generating or replaying I/O traces—captured via Recorder—in real time without modifying application code. FBench supports both POSIX and MPI-IO interfaces and enables configuration-driven exploration through JSON-defined optimization strategies. It faithfully reproduces real-world workloads such as IOR, HACC-IO, FLASH Sedov, and LAMMPS. Evaluations on Lustre reveal that collective I/O write bandwidth can be up to 30× lower than ideal, burst buffers improve non-collective write bandwidth by 1.5×, and performance gains of up to 8× are achievable in LAMMPS scenarios, significantly accelerating I/O optimization studies.

access patternsfile system configurationHPC I/O

Understanding Power and Energy Utilization in Large Scale Production Physics Simulation Codes

Jan 04, 2022
BS
Brian S. Ryujin
🏛️ Lawrence Livermore National Laboratory | NVIDIA | National Nuclear Security Administration US Department of Energy | Sandia National Laboratory

Processor thermal design power (TDP) is widely misused as a proxy for actual power consumption in physics simulations, leading to inaccurate energy-efficiency assessments. Method: This study conducts the first empirical power and energy measurements of major production-scale physics simulation codes on heterogeneous exascale supercomputers at LLNL and Sandia. Leveraging multi-granularity energy modeling, cross-platform benchmarking, and real-time monitoring across commercial and advanced CPU–GPU heterogeneous nodes, it systematically quantifies runtime energy efficiency. Contribution/Results: Under typical simulation workloads, measured power draw is only 30–60% of TDP—substantially lower than nominal ratings. This work challenges the longstanding practice of substituting TDP for measured power, establishing an empirically grounded methodology for evaluating energy efficiency in exascale systems. It provides critical, reproducible, and generalizable energy benchmarks to guide hardware deployment and energy-aware optimization, thereby advancing low-carbon scientific computing.

Compare TDP with actual simulation energy efficiencyEvaluate energy efficiency of advanced computing architecturesMeasure power usage in large-scale physics simulations

Latest Papers

What's happening recently
View more

This work addresses the high overhead of processing massive telemetry data in exascale supercomputing systems by proposing a heterogeneous acceleration–enabled, high-performance diagnostic framework. Integrating high-throughput C++ APIs with GPU-parallelized computation, the framework supports scalable MPI trace analysis and seamless integration with external tools. It introduces a novel topology-aware workflow that maps logical performance anomalies onto the physical coordinates of the Slingshot interconnect and pioneers a three-dimensional performance model to iteratively reconstruct application behavior, enabling precise identification of performance headroom. Evaluated on Aurora, the system ingests traces from 100,000 MPI ranks in just 9.69 seconds, achieving up to a 314× speedup over CPU-based analysis. On Frontier, it uncovers 32.28% potential acceleration for the GAMESS application.

exascalehigh-throughput diagnosticsMPI scalability

Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.

benchmarkingconfiguration spaceLLM infrastructure

This work addresses the absence of a systematic, traceable, and reproducible framework for reporting the performance of mathematical libraries—a gap that hinders accurate performance evaluation and resource planning for scientific applications on high-performance computing (HPC) systems. To this end, the paper introduces LAAB, the first framework explicitly designed around four core principles: traceability, compatibility, reliability, and accessibility. LAAB establishes an end-to-end reproducible performance evaluation pipeline through standardized benchmarking protocols, comprehensive metadata management, execution environment tracking, and advanced performance analysis techniques. The framework substantially enhances the accuracy and interoperability of mathematical library performance reporting, thereby providing a robust foundation for performance prediction and resource scheduling in scientific computing.

benchmarkingHPC systemsmathematical libraries

Hot Scholars

RG

Roberto Grossi

Professor of Computer Science, University of Pisa, Italy
Algorithms and Data Structures
BM

Bongki Moon

Professor of Computer Science & Engineering, Seoul Naitonal University
Database SystemsScalable Systems
GM

Giovanni Manzini

University of Pisa
Algorithms and Data StructuresData CompressionText Searching
TG

Travis Gagie

Associate Professor at Dalhousie University
data structuresdata compression
LD

Loris D'Antoni

University of California-San Diego
Program synthesisProgramming LanguagesProgram AnalysisLLMs for code