🤖 AI Summary
This work addresses the challenge of efficiently parallelizing input-dependent affine recurrences—such as the selective scan in Mamba—on GPUs, which is hindered by their strong sequential dependencies. The paper introduces the first compiler-level abstraction for affine recurrences, leveraging the MLIR framework to automatically transform them into associative Blelloch scans and generate end-to-end optimized GPU code. While preserving the original local recurrence semantics, this approach significantly outperforms sequential baselines implemented in PyTorch and CUDA, achieving performance on par with Mamba’s hand-optimized fused kernels. The proposed method thus enables automatic parallelization and high-performance execution of selective scans without manual kernel engineering.
📝 Abstract
Selective state-space models such as Mamba highlight the practical importance of input-dependent scan recurrences, which preserve linear-time sequence modeling while improving language modeling capabilities. However, these recurrences introduce stricter sequential dependencies than classical structured SSMs, limiting parallel execution on modern accelerators.
We present \textbf{ScanWeaver}, a compiler framework that transforms recurrence-based computations into associative scan representations and lowers them end-to-end to executable GPU programs. We use Mamba-style selective scan as a motivating example of a broader class of affine recurrences that arise in modern ML workloads. Rather than targeting a single model family, ScanWeaver elevates this recurrence structure to a first-class compiler abstraction, enabling systematic MLIR-based lowering to compiler-generated Blelloch scan execution on GPUs.
Across forward selective-scan workloads with matched local recurrence semantics, we validate affine recurrence decomposition, Blelloch lowering, MLIR GPU lowering, executable artifact generation, and actual GPU execution from generated MLIR. We benchmark the resulting ScanWeaver GPU execution against PyTorch and CUDA sequential baselines, and include the Mamba kernel as a fused production baseline for systems context.