roofline performance modeling

Constructs and applies roofline models that map operational intensity (work per byte) to achievable computational throughput by combining hardware peak compute and memory-bandwidth ceilings. Uses measured metrics such as FLOPS and memory bandwidth to analyze kernels or systems, identify whether code is compute- or memory-bound, quantify headroom to theoretical peaks, and guide optimization decisions.

rooflineperformancemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$229K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ridgeline: A 2D Roofline Model for Distributed Systems

Sep 03, 2022
FC
Fabio Checconi
🏛️ Intel

To address the challenge of unified modeling for multi-dimensional bottlenecks—computation, memory, and network—in distributed systems, this paper proposes Ridgeline, the first two-dimensional Roofline performance modeling framework tailored for distributed scenarios. Ridgeline extends the classical Roofline model by incorporating network bandwidth as a core dimension, establishing a dual-axis coordinate system spanned by operational intensity and communication intensity. This enables unified characterization of all three resource constraints and precise identification of the dominant bottleneck. By generalizing Roofline boundary analysis to account for communication overhead, Ridgeline supports communication-aware prediction of multi-node performance ceilings. Evaluated on data-parallel MLP training, it accurately distinguishes communication-bound from compute-bound regimes and successfully predicts performance scaling inflection points across nodes. Ridgeline thus provides a principled, interpretable, and quantifiable theoretical tool for performance diagnosis and optimization in distributed AI systems.

Accounts for network impact on system performanceExtends Roofline model to distributed systemsIdentifies compute, memory, and network bottlenecks

This work addresses the challenge of uniformly evaluating inference efficiency of small language models (SLMs) on resource-constrained edge devices, where objective cross-hardware benchmarks remain scarce. To this end, we propose a systematic evaluation framework grounded in the Roofline model, introducing a novel metric—relative inference potential—that leverages operational intensity (OI) to jointly characterize hardware constraints and model architecture, thereby defining distinct inference potential regions. Empirical analysis reveals how sequence length and model depth influence performance and OI, uncovers efficiency pitfalls arising from hardware heterogeneity, and demonstrates that architectural optimizations such as Multi-Head Latent Attention (MLA) can effectively unlock hardware potential. This study provides both theoretical foundations and practical guidance for hardware-software co-design in edge-side intelligence.

hardware heterogeneityon-device LLMsoperational intensity

How to Keep Pushing ML Accelerator Performance? Know Your Rooflines!

May 22, 2025
MV
Marian Verhelst
🏛️ KU Leuven | imec | ETH Zurich | Universit.a di Bologna | Princeton University | EnCharge AI

To address the performance bottlenecks of machine learning (ML) accelerators under growing model sizes and stringent energy-efficiency constraints, this paper proposes the first enhanced Roofline model deeply co-designed for ML accelerator characteristics. Our method introduces *execution paradigm boundary analysis* and *energy-constrained performance upper-bound modeling*, unifying the quantification of computational intensity, memory hierarchy, and data layout effects on both performance and energy efficiency. By integrating ML workload feature extraction, architecture-level quantitative analysis, and a hardware-algorithm co-evaluation framework, we systematically identify performance bottlenecks across mainstream accelerators for diverse operators and memory layouts. Experimental results reveal synergistic optimization pathways between memory bandwidth and computational density, and clarify several open research directions. The proposed model provides both theoretical foundations and practical guidance for energy-aware ML accelerator architecture design.

Applying roofline model to optimize execution regimesEnhancing ML accelerator performance and efficiencyUnderstanding compute-memory interactions for system efficiency

This work addresses the pronounced performance fluctuations in GEMM operations across adjacent problem sizes—e.g., a 128-element change in dimension N causing up to 30% throughput variation—a phenomenon poorly explained by traditional roofline models and herein termed “performance ruggedness.” The study formally defines and quantifies this effect, modeling GPU performance as a multidimensional surface and distinguishing between software-tunable and hardware-inherent factors. Building on this insight, the authors propose a two-stage runtime optimization strategy combining dynamic tile selection with dynamic-programming-based padding and splitting, achieving O(1) lookup overhead. Evaluated on an Intel Battlemage GPU across 32,768 BF16 GEMM configurations, the approach yields a 30% average throughput improvement and reduces performance ruggedness from 16.8 to approximately 5.0 TFLOPs per 128-step size increment, with residual variations attributed to four hardware-bound sources, thereby delineating the practical limits of software-level optimization.

GEMMGPU performanceperformance ruggedness

Existing performance analysis tools struggle to simultaneously capture temporal dynamics and a holistic view of performance bottlenecks: Roofline models neglect time evolution, while profilers and tracers obscure theoretical performance limits. This work proposes campaign diagrams—a novel visualization framework that uniquely integrates temporal phases with multidimensional resource utilization, including computational throughput, memory bandwidth, data traffic, and latency. Campaign diagrams can be generated from analytical models, simulations, or profiling data, concurrently displaying both theoretical performance ceilings and achieved performance. The approach uncovers cross-phase optimization opportunities often missed by conventional tools, such as counterintuitive cases where enhancing low-intensity operators improves end-to-end performance. Validated on low-rank GEMM and Mamba workloads, the method successfully identifies potential for operator fusion and pipeline optimizations, demonstrating its efficacy in diagnosing deep-rooted performance bottlenecks.

optimization opportunitiesperformance bottlenecksphase-level visualization

Latest Papers

What's happening recently
View more

Existing tools lack automated, cache-aware Roofline modeling support across multiple CPU architectures, hindering effective optimization guidance for high-performance computing applications. This work proposes CARM, the first unified and automated modeling framework spanning x86, ARM, and RISC-V architectures. CARM employs assembly-level microbenchmarks to automatically characterize computational throughput and bandwidth across the entire memory hierarchy, and integrates hardware performance counters with dynamic binary instrumentation for fine-grained bottleneck analysis. The framework supports vectorization across multiple instruction set architectures, achieving a maximum deviation of less than 1% in constructed performance roofs across diverse platforms. By delivering high accuracy and broad applicability, CARM significantly fills the tooling gap for architectures such as AMD and RISC-V in this domain.

automatic benchmarkingCache-aware Roofline Modelcross-architecture

This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.

code representationcomputational kernelsHPC benchmarks

Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.

benchmarkingconfiguration spaceLLM infrastructure

This work addresses the limited feature coverage in high-performance computing caused by hardware performance counters constrained by the number of simultaneously collectible metrics. To overcome this limitation without relying on hardware multiplexing, the authors propose a heuristic multi-run execution trace merging method that aligns and fuses counter data collected across multiple program executions. By analyzing MPI communication structures, timing patterns, and behavioral characteristics, the approach constructs a high-dimensional, unified synthetic trace that expands the effective feature space. This enriched representation enables the training of more comprehensive machine learning–based performance models. Experimental evaluation on the MareNostrum5 platform demonstrates that the merged counters retain high accuracy and significantly improve performance prediction for diverse kernel functions and real-world applications.

execution tracesfeature coveragehardware counters

Hot Scholars

SM

Satoshi Matsuoka

RIKEN Center for Computational Science (R-CCS) / Tokyo Institute of Technology
HPCBig DataScalable AIGreen Computing
MV

Marian Verhelst

Micas - ESAT - KU Leuven, Belgium
Low-energy chip designsensor fusionmachine learningcross-layer optimization
CK

Christos Kozyrakis

Stanford University
Computer ArchitectureComputer SystemsCloud Computing
SD

Sana Damani

NVIDIA
CompilersComputer ArchitectureGPU
SK

Siva Kumar Sastry Hari

Sr. Research Scientist, NVIDIA
Computer ArchitectureAutonomous SystemsHPCReliability