Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited transferability of optimization knowledge in dataflow architectures caused by reliance on vendor libraries. To this end, it proposes Loom, a symbolic compiler that formulates SPMD compilation as hardware-explicit static optimization while maintaining parameters in symbolic form. By jointly optimizing schedules and parameters through CP-SAT constraint solving, legality derivation, and spatial mapping enumeration, Loom establishes the first tuning-free symbolic compilation framework, enabling cross-architecture retargetable and interpretable compiler optimizations. Evaluated on the Tenstorrent architecture, Loom matches or surpasses vendor library performance on workloads such as GEMM without requiring shape-by-shape profiling.
📝 Abstract
Modern AI and HPC accelerators increasingly expose dataflow features: software-visible mechanisms for data movement and overlap, such as inter-core communication through the on-chip network and intra-core asynchronous pipelining. These features shift scheduling responsibility from hardware to the compiler, and because placement, movement, and synchronization become software-visible, they also make the performance of static schedules predictable. Yet high performance on such hardware still relies on vendor-engineered kernel libraries or profile-based auto-tuning, whose embedded expert knowledge transfers poorly across architectures and algorithms. We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures. The central idea is to treat tile-based SPMD compilation as a hardware-explicit static optimization problem. Loom enumerates discrete spatial-mapping and communication candidates while keeping value parameters, such as tiling factors and pipeline knobs, symbolic within each candidate. From an explicit hardware description, it derives symbolic legality constraints and latency expressions, formulates one CP-SAT problem per schedule candidate, and jointly solves inter-core dataflow, intra-core asynchronous scheduling, and block sizes at compile time. On two Tenstorrent generations, Wormhole and Blackhole, Loom matches or exceeds the vendor-optimized TTNN library on GEMM, Flash Attention, and Flash Decode, out of the box and without per-shape profiling or profile-based platform-specific schedule tuning. These results suggest that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.
Problem

Research questions and friction points this paper is trying to address.

dataflow architectures
compiler scheduling
auto-tuning
tile programs
performance portability
Innovation

Methods, ideas, or system contributions that make the work stand out.

symbolic compilation
dataflow architectures
tuning-free
CP-SAT optimization
tile-based SPMD
H
Heru Wang
School of Computing, National University of Singapore, Singapore
W
Wei Li
School of Computing, National University of Singapore, Singapore
Z
Zhenyu Bai
School of Computing, National University of Singapore, Singapore
Tulika Mitra
Tulika Mitra
Professor of Computer Science, National University of Singapore
Design AutomationLow Power DesignEmbedded SystemsReal-Time Systems