empirical hardware characterization

Design and execute empirical tests and microbenchmarks to measure, model, and reverse-engineer hardware and system performance characteristics — including throughput, latency, operand layout and precision, and resource limits. Produce device capability profiles, identify system bottlenecks, and aggregate or cluster devices by observed capabilities to inform modeling, optimization, or system design decisions.

empiricalhardwarecharacterization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An HPC Benchmark Survey and Taxonomy for Characterization

Sep 10, 2025
AH
Andreas Herten
🏛️ Forschungszentrum Jülich | Lawrence Livermore National Laboratory | Texas A&M University

The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.

Developing taxonomy to categorize HPC benchmarks systematicallyProviding structured comparison of hardware and software evaluation toolsSurveying existing HPC benchmarks for comprehensive characterization

This study systematically evaluates the reliability of four machine learning–based ranking models—NeuroScalar, SimNet, Concorde, and OneDSE—in ordering hardware configurations at the program phase level for microarchitectural design space exploration. Across structural parameter and behavioral policy scenarios, the analysis—integrating cycle-accurate simulation, Bayesian accuracy assessment, and information-theoretic methods—reveals, for the first time, that a substantial fraction (22.4%) of program windows exhibit counterintuitive rankings in structural settings, and inter-model consistency remains low (23.3%–39.9%). In behavioral policy scenarios, most models fail to surpass a featureless baseline, with the best achieving only a 2.1-percentage-point improvement. The work further establishes a theoretical upper bound on ranking accuracy when critical microarchitectural states are unobservable, demonstrating inherent limitations of instruction-stream–based approaches.

cycle-level simulationdesign-space explorationmachine-learned ranking

$mu$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics

Jan 22, 2025
NJ
Nils Japke
🏛️ TU Berlin | Simula Research Laboratory

To address the low efficiency of performance monitoring in CI/CD pipelines caused by excessive repetition in microbenchmarking, this paper proposes a static, data-driven method for determining the minimal reliable repetition count per microbenchmark. Leveraging historical execution data and five statistical stability metrics—including coefficient of variation (CV), relative standard deviation (RSD), and interquartile range (IQR)—our approach statically infers the minimum repetitions required to achieve measurement reliability. It explicitly models JVM warm-up effects and supports cross-language adaptation for both Java and Go. Furthermore, the method is tightly integrated with industrial-grade performance regression detection frameworks. Experimental evaluation across 14 open-source projects demonstrates that, while preserving regression detection accuracy, our technique reduces microbenchmark measurement time by 95.83% (Go) and 94.17% (Java) on average.

Efficiency in TestingPerformance MonitoringSoftware Development

This work addresses the growing mismatch between modern datacenter and AI workloads and traditional CPU benchmarks, which often fail to accurately capture performance bottlenecks. The study presents the first systematic microarchitectural analysis of SPEC CPU2026 across nine mainstream processors, leveraging cross-platform performance counters, clustering algorithms, and detailed case studies—such as page size effects and prefetching strategies—to construct a highly representative compact subset comprising only four to five programs. This subset preserves 96.4%–99.9% of the original benchmark suite’s behavioral characteristics. Furthermore, the authors introduce a polling-interleaved pattern to synthesize proxy workloads, reducing the IPC gap with real-world DCPerf workloads to just 13.7%, thereby significantly enhancing both the efficiency and representativeness of CPU performance evaluation.

CPU benchmarkingcross-suite comparisonmicroarchitectural bottlenecks

Latest Papers

What's happening recently
View more

This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.

cluster-wide benchmarkingcontinuous testingdata center validation

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

This work addresses the inefficiency of traditional approaches to predicting workload performance under varying memory configurations, which typically rely on time-consuming simulations or repeated measurements. The study reveals, for the first time, a predictable relationship between cycles per instruction (CPI) and maximum memory stall across diverse workloads. By leveraging hardware performance counters collected from a single native execution—combined with mechanistic insights and empirical data—the authors construct a regression model that enables highly accurate, simulation-free first-order performance prediction. Evaluated across six machine configurations and two simulators, the method reduces CPI prediction error by 2× compared to the best existing single-run techniques. On ARM servers, it achieves a median error of 12.7% and a 90th-percentile error of 35.9%, maintaining robust accuracy even when extrapolating to memory latencies up to 8× higher than baseline.

CPI modelingmemory stallone-shot prediction

This study addresses the accuracy bottleneck in machine shape prediction caused by missing critical metrics such as memory and power consumption. We propose a heterogeneous device cluster modeling approach that refines the machine shape model by indexing five unmeasured fields and introducing hardware reachability Boolean values. To eliminate interference from heuristic tuning, deterministic compiler directive standards are established, and a weighted communication-free partitioning algorithm is designed for efficient solving. Extensive evaluations across Intel, AMD, and NVIDIA platforms demonstrate that this method successfully closes all prediction fields on five devices while quantifying reservation overheads. Furthermore, it reveals potential adverse effects of excluded heuristics and achieves precise alignment between predictions and actual hardware behavior.

hardware mappingheterogeneous computingmachine shape

Hot Scholars

AO

Ataberk Olgun

ETH Zurich
Computer ArchitectureMemory SystemsComputer SecurityReliability
YF

Yunsi Fei

Professor of Electrical and Computer Engineering, Northeastern University
hardware securityEDAcomputer architectureembedded systems
SY

Shimeng Yu

Georgia Institute of Technology, Dean's Professor
Non-volatile MemoryRRAMFerroelectric MemoriesIn-Memory Computing
IK

Ian Karlin

Lawrence Livermore National Laboratory
Y(

Yingyan (Celine) Lin

Associate Professor, Georgia Institute of Technology
Efficient AI algorithmsDeep learning acceleratorsGreen AI