TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of systematic comparisons between tile-based programming models such as Triton and cuTile by constructing a controlled benchmark for standardized evaluation on NVIDIA B200 GPUs. Methodologically, we design a unified test suite comprising 45 operators and integrate PyTorch reference implementations, auto-tuning, Roofline modeling, and GPU profiling techniques for comprehensive bottleneck diagnosis. Furthermore, this work presents the first quantitative comparison of token efficiency for LLM-generated code. Our findings reveal that cuTile excels in Tensor Core-intensive kernels, whereas Triton demonstrates superior performance and greater token efficiency in irregular and memory-bandwidth-bound scenarios.
📝 Abstract
Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.
Problem

Research questions and friction points this paper is trying to address.

Tile-based programming models
performance evaluation
bottleneck diagnosis
benchmark
kernel development
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tile-based programming models
Benchmark
Performance evaluation
Bottleneck diagnosis
LLM-generated kernels
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bowen Cui
George Mason University
Z
Zhongchun Zhou
George Mason University
H
Hao Wu
George Mason University
T
Tejas Ramesh
George Mason University
J
Junyu Yin
George Mason University
J
Jialiang Gu
George Mason University
Keren Zhou
Keren Zhou
George Mason University
Concurrent ProgrammingDistributed SystemParallel ProgrammingMachine Learning