🤖 AI Summary
Sparse matrix-matrix multiplication (SPMM) on GPUs suffers from irregular memory access patterns, resulting in low cache utilization and imbalanced compute unit workload. To address this, we propose a compiler-level transformation—“enumerate-and-sparse-coarsen”—that jointly optimizes sparse-pattern-aware tiling, register-level data layout rearrangement, and workload coarsening for the first time. This holistic approach enhances data reuse in registers and caches while ensuring balanced occupancy across streaming multiprocessors (SMs), all without manual tuning and fully automated by the compiler. Evaluated on an NVIDIA A100 GPU, our method achieves a 1.84×–2.27× geometric mean speedup over cuBLAS/cuSPARSE for sparse neural network workloads—including sparse convolutional and Transformer models—significantly improving end-to-end inference efficiency of sparse deep neural networks on GPUs.
📝 Abstract
Sparse data structures are commonly used in neural networks to reduce the memory footprint. These data structures are compact but cause irregularities such as random memory accesses, which prevent efficient use of the memory hierarchy. GPUs are a common platform for machine learning practitioners, but running compact data structures on these devices often leads to slow-downs due to inefficient use of computing and memory resources. This paper proposes a new compiler transformation, enumerate-and-sparse-coarsen, that accelerates sparse matrix-matrix multiplication (SPMM) on GPU devices. The transformation increases data reuse in registers and caches while creating more balanced workloads for GPU computing resources. The transformation is tested on sparse neural networks in convolutional and transformer models. On an A100 GPU and across a columns of matrix B (bCols) in $ A imes B = C$ from range of 32 to 128, the transformation yields a geometric mean speedup of 1.84$ imes$ to 2.27$ imes$ compared to cuBLAS and cuSPARSE baselines, respectively.