Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models

📅 2026-04-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently deploying Mamba models on modern hardware, which is hindered by complex data dependencies leading to excessive off-chip memory accesses. To tackle this, the authors introduce a cascade-of-Einsums abstraction to systematically characterize Mamba’s computational structure and devise cross-Einsum fusion strategies that minimize data movement. Building upon this foundation, they design Mambalaya, a reconfigurable accelerator that pioneers the use of an extended Einsum framework to uniformly fuse multiple operators within Mamba. Experimental results demonstrate that Mambalaya achieves a 4.9× speedup over MARCA during the prefill phase and a 1.9× speedup in the generation phase; moreover, in prefill-dominated scenarios, it delivers 1.5× higher performance than existing fused accelerators.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Learning to SearchNatural Language Processing: Code Generation / Program Synthesis from Natural Language

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Mamba is an emerging, complex workload with various short-range and long-range dependencies, nonlinearities, and elementwise computations that are unable to run at near-peak speeds on modern hardware. Specifically, Mamba's complex dependency graph makes fusion across its full operator cascade difficult, leaving substantial inter-operator memory traffic on the table. To address these challenges, we propose Mambalaya, a novel reconfigurable accelerator that leverages fusion to overcome the limitations of Mamba. We use the recently proposed cascade-of-Einsums abstraction to characterize Mamba's full computational structure, then apply the extended Einsum framework to systematically explore inter-Einsum fusion opportunities. This principled approach yields a series of fusion mappings that reduce off-chip inter-Einsum traffic. These mappings are supported by the underlying Mambalaya architecture. Mambalaya achieves a layer performance speedup of 4.9$\times$ for prefill and 1.9$\times$ for generation over MARCA. In prefill-dominated scenarios, it achieves up to 1.5$\times$ over a recent fine-grained, memory-aware fusion accelerator for Mamba.
Problem

Research questions and friction points this paper is trying to address.

Mamba
fusion
state-space models
memory traffic
Einsum
Innovation

Methods, ideas, or system contributions that make the work stand out.

Einsum-based fusion
State-Space Models
Reconfigurable accelerator
Memory traffic reduction
Operator fusion
T
Toluwanimi O. Odemuyiwa
University of California, Davis
J
John D. Owens
University of California, Davis
J
Joel S. Emer
MIT / NVIDIA
M
Michael Pellauer
NVIDIA