Score
Designs and implements compiler pipelines that decompose rendering workloads into statically-shaped, parallel primitives and represent those primitives with static shapes suitable for ahead-of-time compilation. Builds cross-platform lowering passes and XLA-backed code generation that transform those primitives into device-specific executable code for GPUs, TPUs, or CPUs, producing optimized kernels without hand-writing hardware-specific implementations.
This study addresses the lack of a unified instruction set architecture across GPU vendors, which hinders efficient cross-platform portability of parallel programs. Through a systematic analysis of instruction sets from sixteen microarchitectures spanning four major vendors, the work identifies ten cross-platform computational primitives, six dialect-like variations, and six fundamental architectural divergences. Leveraging these insights, it proposes the first vendor-agnostic abstract execution model for GPUs. Validated against official documentation, patents, reverse-engineered data, and cross-platform benchmarks, the model demonstrates strong performance on architecturally disparate hardware—specifically NVIDIA T4 and Apple M1—matching or exceeding native performance in five out of six benchmark suites, with only parallel reduction lagging at 62.5% efficiency, thereby underscoring the critical role of the shuffle primitive.
Large language model (LLM) inference suffers from poor performance portability across heterogeneous GPUs (e.g., NVIDIA, AMD, Intel), heavy reliance on vendor-provided closed-source optimizations, and labor-intensive manual kernel tuning. Method: We propose a portable, high-performance execution framework that requires no user code modification, achieved by tightly integrating just-in-time (JIT) compilation with fine-grained kernel parameter autotuning—enabling joint compile-time optimization between JIT and autotuning for the first time. This synergy expands the configuration search space by 15× and significantly enhances generated kernel diversity. Results: Evaluated on Flash Attention, our approach outperforms vendor-optimized libraries across all three GPU architectures, achieving up to 230% speedup. Generated kernels are 70× smaller in binary size, and manual tuning is fully eliminated. Our method establishes a new paradigm for efficient, cross-platform LLM deployment.
CPU-GPU data transfers in AI-based code generation impose severe compilation and iterative latency bottlenecks. Method: This paper introduces the first theoretical framework for GPU-native compilation, systematically proposing and analyzing three paradigms—parallel traditional compilation, neural compilation, and hybrid compilation. Contributions/Results: (1) A probabilistic formal verification mechanism that jointly optimizes accuracy and parallelism; (2) Theoretical characterization of latency, energy, and correctness trade-offs across all three paradigms, along with provable co-optimization pathways; (3) A deployable hybrid architecture ensuring strict correctness. Experimental and theoretical analysis demonstrates up to 10–100× speedup in end-to-end code iteration latency: traditional GPU compilation accelerates by 2–5×, neural compilation by 10–100×, and the hybrid approach achieves practical efficiency without compromising formal correctness.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
Modern GPU SIMT programming models exhibit a semantic gap with underlying hardware task-parallelism, forcing developers to manually implement warp-level specialization and inter-warp communication—resulting in high development overhead and error-prone code. This paper proposes a compiler-driven approach to automated warp specialization. We introduce *asynchronous references* (arefs) as an intermediate representation that uniformly models inter-warp data dependencies and asynchronous communication. Leveraging high-level tiling annotations, our compiler automatically infers producer-consumer roles without modifying kernel source code. By integrating dataflow pipeline optimization with hardware-aware scheduling, we generate high-performance LLM operator kernels for NVIDIA H100. Experiments show our generated GEMM achieves 2.1× the performance of cuBLAS; our attention kernel outperforms Triton by 1.2×; and both match the performance of hand-tuned CUTLASS kernels.
Modern GPU programming faces a fundamental trade-off between abstraction level and hardware control: excessive abstraction impedes performance optimization, while overly low-level approaches impose significant development burdens. This work proposes TLX, an extension to the Triton language based on a Multi-Instruction, Multi-Warp (MIMW) execution model—the first to integrate MIMW into a high-level GPU programming framework. TLX operates at the warp-group granularity, explicitly supporting multi-warp scheduling, shared memory orchestration, asynchronous operations, and cluster-aware control flow. It preserves Triton’s elegant block-level programming model while enabling efficient exploitation of native hardware features. Experimental results demonstrate that TLX kernels achieve state-of-the-art performance with substantially reduced development effort and have been successfully deployed in large-scale training and inference systems.
This work proposes a novel approach within the BuildIt system to circumvent the engineering complexity and high cost of traditional program optimization, which relies on backward data-flow analysis to obtain future execution information. Instead of performing backward analysis, the method introduces oracle variables that enable multiple forward executions to predict future program behavior. This is the first application of oracle variables to support optimizations requiring future information, eliminating the need for intermediate representations or custom parsers. Consequently, it significantly reduces implementation overhead for embedded domain-specific languages (DSLs) built atop C++. Integrated with staged compilation and code generation targeting C, C++, and CUDA, experimental results demonstrate that the approach not only enhances performance but also substantially decreases the engineering effort required for DSL development.
This work addresses the challenge of efficiently generating vector-length-agnostic (VLA) machine learning code for scalable vector instruction sets such as Arm SVE, where unknown vector lengths at compile time hinder traditional compilers. The authors present the first end-to-end VLA support in MLIR/IREE, introducing a vector-length-aware compact data layout and unifying dynamic tiling, operator fusion, and scalable vectorization within a single compilation framework. Evaluated on Arm CPUs, the generated SVE code achieves up to 1.45× speedup over IREE’s NEON implementation, outperforms multiple frameworks in the PyTorch ecosystem, and demonstrates strong scalability with increasing vector lengths in simulation, effectively balancing performance and hardware portability.
This work addresses the challenge that lightweight compilers and source-to-source tools struggle to reuse the sophisticated inlining heuristics of mature compilers like GCC or LLVM due to their reliance on complex intermediate representations and analysis infrastructures. To bridge this gap, the paper introduces the first portable inlining prediction framework that leverages diagnostic outputs from production compilers as supervision signals. By extracting call-site features through AST normalization and constructing a lightweight structured IR, the approach trains tabular models—such as CatBoost—that can be directly compiled into pure C code without runtime dependencies. Evaluated on a dataset of 330,000 call sites, the model achieves a ROC-AUC of 0.928 and PR-AUC of 0.713; with threshold tuning, it attains an F1 score of 0.729 while reducing the false positive rate to 0.084, thereby enabling the first practical transfer of industrial-grade inlining decisions to resource-constrained systems.