A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs

📅 2025-06-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Sparse matrix-matrix multiplication (SPMM) on GPUs suffers from irregular memory access patterns, resulting in low cache utilization and imbalanced compute unit workload. To address this, we propose a compiler-level transformation—“enumerate-and-sparse-coarsen”—that jointly optimizes sparse-pattern-aware tiling, register-level data layout rearrangement, and workload coarsening for the first time. This holistic approach enhances data reuse in registers and caches while ensuring balanced occupancy across streaming multiprocessors (SMs), all without manual tuning and fully automated by the compiler. Evaluated on an NVIDIA A100 GPU, our method achieves a 1.84×–2.27× geometric mean speedup over cuBLAS/cuSPARSE for sparse neural network workloads—including sparse convolutional and Transformer models—significantly improving end-to-end inference efficiency of sparse deep neural networks on GPUs.

Technology Category

Machine Learning: Hardware-aware MLData Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Sparse data structures are commonly used in neural networks to reduce the memory footprint. These data structures are compact but cause irregularities such as random memory accesses, which prevent efficient use of the memory hierarchy. GPUs are a common platform for machine learning practitioners, but running compact data structures on these devices often leads to slow-downs due to inefficient use of computing and memory resources. This paper proposes a new compiler transformation, enumerate-and-sparse-coarsen, that accelerates sparse matrix-matrix multiplication (SPMM) on GPU devices. The transformation increases data reuse in registers and caches while creating more balanced workloads for GPU computing resources. The transformation is tested on sparse neural networks in convolutional and transformer models. On an A100 GPU and across a columns of matrix B (bCols) in $ A imes B = C$ from range of 32 to 128, the transformation yields a geometric mean speedup of 1.84$ imes$ to 2.27$ imes$ compared to cuBLAS and cuSPARSE baselines, respectively.
Problem

Research questions and friction points this paper is trying to address.

Accelerates sparse matrix multiplication on GPUs
Reduces memory access irregularities in sparse data
Improves workload balance and data reuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compiler transformation for sparse matrix multiplication
Enhances data reuse in registers and caches
Balances GPU workloads for faster execution
🔎 Similar Papers
No similar papers found.