Annotation-Driven Migration of CUDA Programs to Tenstorrent Blackhole

๐Ÿ“… 2026-10-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge that CUDAโ€™s implicit scheduling model is ill-suited for Blackholeโ€™s explicit distributed architecture, which impedes the migration of high-performance computing (HPC) kernels. To overcome this, we propose an MLIR-based, annotation-driven compilation framework. By leveraging static affine analysis and declarative annotations, the method automatically derives data placement, inter-core communication patterns, and Tensix compute mappings. Furthermore, it introduces a novel spatial mapping mechanism that uncovers nonlinear coupling relationships among strategies, enabling globally coordinated configuration optimization. Experimental evaluations demonstrate that the proposed approach achieves 4.2ร— and 2.1โ€“2.2ร— speedups for BF16 and FP32 Gaussian elimination, respectively. These results validate the feasibility and effectiveness of the framework in efficiently migrating HPC kernels to explicitly scheduled architectures.
๐Ÿ“ Abstract
Tenstorrent Blackhole combines distributed local memories, explicit inter-core communication, and Tensix cores decoupling data movement from computation. CUDA offers a substantial HPC software base but leaves physical data placement and scheduling largely implicit. Migrating CUDA HPC kernels to Blackhole requires spatial mapping across cores and per-core coordination of compute and data-movement kernels. We present an MLIR-based compiler deriving data placement, inter-core communication, and tile computation from a statically shaped affine CUDA subset. Declarative annotations express choices not fixed by the source, including compute/DM operation placement and streaming granularity. We evaluate it on Gaussian elimination, five-point stencil, and symmetric rank-k update. Relative to our compiler's default realizations, the best measured configurations achieve a 4.2x speedup for BF16 Gaussian and 2.1-2.2x for FP32 Gaussian and Jacobi. A choice's performance impact can reverse with surrounding policies, precision, and loop schedule, motivating comparison of alternative per-core realizations rather than independent policy selection.
Problem

Research questions and friction points this paper is trying to address.

CUDA migration
Tenstorrent Blackhole
spatial mapping
data placement
inter-core communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

MLIR-based compiler
CUDA migration
declarative annotations
data placement
Tenstorrent Blackhole
๐Ÿ”Ž Similar Papers
2024-06-302024 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops)Citations: 3
2024-02-14Proceedings of the 39th ACM International Conference on SupercomputingCitations: 3