Score
Constructs and applies roofline models that map operational intensity (work per byte) to achievable computational throughput by combining hardware peak compute and memory-bandwidth ceilings. Uses measured metrics such as FLOPS and memory bandwidth to analyze kernels or systems, identify whether code is compute- or memory-bound, quantify headroom to theoretical peaks, and guide optimization decisions.
To address the challenge of unified modeling for multi-dimensional bottlenecks—computation, memory, and network—in distributed systems, this paper proposes Ridgeline, the first two-dimensional Roofline performance modeling framework tailored for distributed scenarios. Ridgeline extends the classical Roofline model by incorporating network bandwidth as a core dimension, establishing a dual-axis coordinate system spanned by operational intensity and communication intensity. This enables unified characterization of all three resource constraints and precise identification of the dominant bottleneck. By generalizing Roofline boundary analysis to account for communication overhead, Ridgeline supports communication-aware prediction of multi-node performance ceilings. Evaluated on data-parallel MLP training, it accurately distinguishes communication-bound from compute-bound regimes and successfully predicts performance scaling inflection points across nodes. Ridgeline thus provides a principled, interpretable, and quantifiable theoretical tool for performance diagnosis and optimization in distributed AI systems.
This work addresses the challenge of uniformly evaluating inference efficiency of small language models (SLMs) on resource-constrained edge devices, where objective cross-hardware benchmarks remain scarce. To this end, we propose a systematic evaluation framework grounded in the Roofline model, introducing a novel metric—relative inference potential—that leverages operational intensity (OI) to jointly characterize hardware constraints and model architecture, thereby defining distinct inference potential regions. Empirical analysis reveals how sequence length and model depth influence performance and OI, uncovers efficiency pitfalls arising from hardware heterogeneity, and demonstrates that architectural optimizations such as Multi-Head Latent Attention (MLA) can effectively unlock hardware potential. This study provides both theoretical foundations and practical guidance for hardware-software co-design in edge-side intelligence.
To address the performance bottlenecks of machine learning (ML) accelerators under growing model sizes and stringent energy-efficiency constraints, this paper proposes the first enhanced Roofline model deeply co-designed for ML accelerator characteristics. Our method introduces *execution paradigm boundary analysis* and *energy-constrained performance upper-bound modeling*, unifying the quantification of computational intensity, memory hierarchy, and data layout effects on both performance and energy efficiency. By integrating ML workload feature extraction, architecture-level quantitative analysis, and a hardware-algorithm co-evaluation framework, we systematically identify performance bottlenecks across mainstream accelerators for diverse operators and memory layouts. Experimental results reveal synergistic optimization pathways between memory bandwidth and computational density, and clarify several open research directions. The proposed model provides both theoretical foundations and practical guidance for energy-aware ML accelerator architecture design.
This work addresses the pronounced performance fluctuations in GEMM operations across adjacent problem sizes—e.g., a 128-element change in dimension N causing up to 30% throughput variation—a phenomenon poorly explained by traditional roofline models and herein termed “performance ruggedness.” The study formally defines and quantifies this effect, modeling GPU performance as a multidimensional surface and distinguishing between software-tunable and hardware-inherent factors. Building on this insight, the authors propose a two-stage runtime optimization strategy combining dynamic tile selection with dynamic-programming-based padding and splitting, achieving O(1) lookup overhead. Evaluated on an Intel Battlemage GPU across 32,768 BF16 GEMM configurations, the approach yields a 30% average throughput improvement and reduces performance ruggedness from 16.8 to approximately 5.0 TFLOPs per 128-step size increment, with residual variations attributed to four hardware-bound sources, thereby delineating the practical limits of software-level optimization.
Existing performance analysis tools struggle to simultaneously capture temporal dynamics and a holistic view of performance bottlenecks: Roofline models neglect time evolution, while profilers and tracers obscure theoretical performance limits. This work proposes campaign diagrams—a novel visualization framework that uniquely integrates temporal phases with multidimensional resource utilization, including computational throughput, memory bandwidth, data traffic, and latency. Campaign diagrams can be generated from analytical models, simulations, or profiling data, concurrently displaying both theoretical performance ceilings and achieved performance. The approach uncovers cross-phase optimization opportunities often missed by conventional tools, such as counterintuitive cases where enhancing low-intensity operators improves end-to-end performance. Validated on low-rank GEMM and Mamba workloads, the method successfully identifies potential for operator fusion and pipeline optimizations, demonstrating its efficacy in diagnosing deep-rooted performance bottlenecks.
Existing tools lack automated, cache-aware Roofline modeling support across multiple CPU architectures, hindering effective optimization guidance for high-performance computing applications. This work proposes CARM, the first unified and automated modeling framework spanning x86, ARM, and RISC-V architectures. CARM employs assembly-level microbenchmarks to automatically characterize computational throughput and bandwidth across the entire memory hierarchy, and integrates hardware performance counters with dynamic binary instrumentation for fine-grained bottleneck analysis. The framework supports vectorization across multiple instruction set architectures, achieving a maximum deviation of less than 1% in constructed performance roofs across diverse platforms. By delivering high accuracy and broad applicability, CARM significantly fills the tooling gap for architectures such as AMD and RISC-V in this domain.
This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.
研究通过结合静态代码特征、动态执行轨迹和内核级资源数据,发现13种性能原型,并提出一个多信号回归检测框架,以提高软件性能分析和预测的准确性。
Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.
This work addresses the limited feature coverage in high-performance computing caused by hardware performance counters constrained by the number of simultaneously collectible metrics. To overcome this limitation without relying on hardware multiplexing, the authors propose a heuristic multi-run execution trace merging method that aligns and fuses counter data collected across multiple program executions. By analyzing MPI communication structures, timing patterns, and behavioral characteristics, the approach constructs a high-dimensional, unified synthetic trace that expands the effective feature space. This enriched representation enables the training of more comprehensive machine learning–based performance models. Experimental evaluation on the MareNostrum5 platform demonstrates that the merged counters retain high accuracy and significantly improve performance prediction for diverse kernel functions and real-world applications.