Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of inefficient programming models for custom AI accelerators, which struggle to support diverse machine learning operators effectively. For the first time, the authors successfully deploy the Triton language on Meta’s in-house MTIA-2i accelerator by introducing a new compiler backend, enhancing TorchInductor’s code generation, and incorporating minimal language extensions tailored to the hardware’s characteristics. This approach bridges the programming gap between high-level ML frameworks and heterogeneous hardware, achieving kernel performance comparable to hand-optimized C++ implementations while preserving developer productivity. The solution has been deployed across approximately 60 model types, covering 50% of network layers and 47% of non-GEMM execution time.
📝 Abstract
The rapid growth in machine learning workloads has fueled the proliferation of custom accelerator architectures. Designed from the ground up, these accelerators often expose programming models that are distinct from GPUs. While hyperscalers and AI chip startups continue to innovate in this space, achieving broad operator coverage to support diverse models remains a major challenge. Additionally, an easy-to-use, high-level kernel programming language is important for rapid iteration of models and kernels. Triton, together with TorchInductor, addresses these issues on GPUs, but its viability on accelerators with different programming models has yet to be established. In this work, we present the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta. To support MTIA-2i, we develop a new compiler backend that targets it, introduce enhancements to TorchInductor code generation, and propose minimal language extensions that expose MTIA-specific architectural features. We demonstrate that Triton-MTIA kernels achieve performance competitive with expert-tuned C++ implementations. Leveraging these development efficiency gains, we successfully deployed manually written and Inductor-generated Triton kernels in production across approximately 60 different model types, accounting for 50% of layers and 47% of non-GEMM execution time for these models. Our results provide compelling evidence that DSLs like Triton can bridge the programming model gaps between ML frameworks, kernels, and custom accelerators, enabling rapid innovation and efficient deployment at scale.
Problem

Research questions and friction points this paper is trying to address.

custom AI accelerators
programming model gaps
operator coverage
kernel programming language
machine learning workloads
Innovation

Methods, ideas, or system contributions that make the work stand out.

Triton
custom AI accelerator
compiler backend
TorchInductor
domain-specific language
🔎 Similar Papers
No similar papers found.
Haishan Zhu
Haishan Zhu
Microsoft
Computer ArchitectureMachine Learning
D
Domi Yan
Meta Platforms
M
Michael Levesque-Dion
Meta Platforms
C
Changxu Zhang
Meta Platforms
M
Mitch Gamburg
Meta Platforms
K
Kirsten Lee
Meta Platforms
G
Giancarlo Colmenares
Meta Platforms
A
Aditya Bhagwat
Meta Platforms
A
Arnab De
Meta Platforms
M
Markus Le Roux
Meta Platforms
V
Victor Perez Carrasco
Meta Platforms
X
Xin Tong
Meta Platforms
W
Will Cromar
Meta Platforms
S
Simran Barnwal
Meta Platforms
A
Andrew Uderian
Meta Platforms
B
Blaine Burton Rister
Meta Platforms
J
Jordan Fix
Meta Platforms
J
Jazlyn Li
Meta Platforms
Z
Zejun Huang
Meta Platforms
L
Lite Ye
Meta Platforms
Nan Zhang
Nan Zhang
Facebook Inc.
Mobile SecurityIoT SecuritySystem SecurityPrivacy
X
Xinchen Guo
Meta Platforms
A
Andiry Xu
Meta Platforms
M
Michael Roberts
Meta Platforms
K
Kunming Ho
Meta Platforms