AI as a Compiler: Compiling Triton kernels without the Triton compiler

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high maintenance costs of compiler backends and the difficulty of adapting to emerging hardware by proposing the AI Lowering paradigm. This approach leverages large language model (LLM) agents to directly compile Triton kernels into PTX code, thereby bypassing conventional optimization pipelines. Furthermore, it substantially extends the Volta validator to accommodate Blackwell architecture features. Experimental results demonstrate that the proposed method achieves performance improvements ranging from 0.83Γ— to 3.34Γ— across diverse GPU platforms, significantly reducing the software deployment overhead associated with new chip architectures.
πŸ“ Abstract
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.
Problem

Research questions and friction points this paper is trying to address.

AI compiler
Triton kernels
PTX lowering
large language models
compiler backend
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Lowering
LLM Compiler
Triton to PTX
Agentic Harness
PTX Verification
πŸ”Ž Similar Papers
No similar papers found.