๐ค AI Summary
This study addresses the challenge that CUDAโs implicit scheduling model is ill-suited for Blackholeโs explicit distributed architecture, which impedes the migration of high-performance computing (HPC) kernels. To overcome this, we propose an MLIR-based, annotation-driven compilation framework. By leveraging static affine analysis and declarative annotations, the method automatically derives data placement, inter-core communication patterns, and Tensix compute mappings. Furthermore, it introduces a novel spatial mapping mechanism that uncovers nonlinear coupling relationships among strategies, enabling globally coordinated configuration optimization. Experimental evaluations demonstrate that the proposed approach achieves 4.2ร and 2.1โ2.2ร speedups for BF16 and FP32 Gaussian elimination, respectively. These results validate the feasibility and effectiveness of the framework in efficiently migrating HPC kernels to explicitly scheduled architectures.
๐ Abstract
Tenstorrent Blackhole combines distributed local memories, explicit inter-core communication, and Tensix cores decoupling data movement from computation. CUDA offers a substantial HPC software base but leaves physical data placement and scheduling largely implicit. Migrating CUDA HPC kernels to Blackhole requires spatial mapping across cores and per-core coordination of compute and data-movement kernels. We present an MLIR-based compiler deriving data placement, inter-core communication, and tile computation from a statically shaped affine CUDA subset. Declarative annotations express choices not fixed by the source, including compute/DM operation placement and streaming granularity. We evaluate it on Gaussian elimination, five-point stencil, and symmetric rank-k update. Relative to our compiler's default realizations, the best measured configurations achieve a 4.2x speedup for BF16 Gaussian and 2.1-2.2x for FP32 Gaussian and Jacobi. A choice's performance impact can reverse with surrounding policies, precision, and loop schedule, motivating comparison of alternative per-core realizations rather than independent policy selection.