From Roofline to Ruggedness: Decomposing and Smoothing the GEMM Performance Landscape

📅 2026-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the pronounced performance fluctuations in GEMM operations across adjacent problem sizes—e.g., a 128-element change in dimension N causing up to 30% throughput variation—a phenomenon poorly explained by traditional roofline models and herein termed “performance ruggedness.” The study formally defines and quantifies this effect, modeling GPU performance as a multidimensional surface and distinguishing between software-tunable and hardware-inherent factors. Building on this insight, the authors propose a two-stage runtime optimization strategy combining dynamic tile selection with dynamic-programming-based padding and splitting, achieving O(1) lookup overhead. Evaluated on an Intel Battlemage GPU across 32,768 BF16 GEMM configurations, the approach yields a 30% average throughput improvement and reduces performance ruggedness from 16.8 to approximately 5.0 TFLOPs per 128-step size increment, with residual variations attributed to four hardware-bound sources, thereby delineating the practical limits of software-level optimization.
📝 Abstract
Adjacent GEMM problems that differ by a single 128-element step in N can show 30% different throughput on the same GPU. This pervasive performance ruggedness - invisible to roofline analysis and peak-FLOPs intuition, yet dominant for every non-peak workload - is the subject of this paper. We propose performance ruggedness analysis as an analytical framework complementary to roofline: rather than summarizing GPU performance with a scalar bound, treat the full multidimensional performance surface as the object of study, decompose its texture into mechanism-attributable components and separate software-removable contributions from hardware-bound ones. The framing is directly analogous to deep-learning loss landscapes - a continuous quantity (the idealized time 2MNK / compute_throughput_peak) made rugged by interaction with discrete hardware substrates (tiles, sub-groups, cache lines, DRAM channels). We apply the framework to BF16 NN (no transpose) GEMM on Intel Battlemage (Arc B580, sycl-tla) via a 32,768-configuration sweep (M, N, K) belongs to {128, ..., 4096}^3. The peak is 110.8 TFLOPs at the non-square shape M=3840, N=2048, K=4096 with the default tile size; the initial landscape roughness is 16.8 TFLOPs per 128-step against an ideal of 2.0. A two-stage software stack - (i) best-of-six dynamic tile selection and (ii) a novel dynamic-programming based padding-and-splitting optimizer with O(1) runtime lookup - reduces roughness by 70% and raises mean throughput by 30%. Cross-tile experiments establish that the residual sawtooth period scales exactly with software tile size, ruling out cache set conflicts and attributing the remaining variance to four hardware-bound sources (per-kernel base overhead, wave quantization, DPAS atom geometry and GDDR6 channel-hash interactions).
Problem

Research questions and friction points this paper is trying to address.

GEMM
performance ruggedness
roofline model
GPU performance
throughput variability
Innovation

Methods, ideas, or system contributions that make the work stand out.

performance ruggedness
GEMM optimization
roofline model
dynamic tiling
hardware-software co-analysis
💼 Related Jobs
No related jobs found.
A
Aditya Chatterjee
Intel Corporation