Score
Implements compilers for machine-learning models by designing and building the pipeline that lowers high-level model graphs into optimized intermediate representations and low-level code, including passes for operator fusion, scheduling, memory/layout transforms, and other optimizations. Produces target-specific code generators and runtime integration (XLA-/TVM-style) to emit efficient kernels and binaries for CPUs, GPUs, and accelerators.
Integrating custom hardware accelerators—particularly GEMM-based ones—into mainstream ML compilers remains challenging due to tight coupling between accelerator-specific optimizations and compiler internals. Method: This paper proposes a high-level, low-intrusion integration methodology for TVM that abstracts hardware scheduling interfaces to decouple accelerator characteristics from compiler implementation. Leveraging the CoSA design-space exploration framework, it automates hardware-aware scheduling optimizations—including tensor tiling, non-uniform mapping, and double buffering—without modifying TVM’s core infrastructure. Contribution/Results: Evaluated on the Gemmini accelerator, the approach achieves performance on par with hand-optimized toolchains while significantly improving developer productivity and cross-model/cross-architecture portability. The abstraction enables seamless reuse of scheduling policies across diverse accelerator microarchitectures and neural network workloads, reducing integration effort from weeks to hours.
This work addresses the challenges posed by dynamic AI models—whose varying tensor shapes and control flows hinder existing compilers from simultaneously achieving fast compilation, low memory overhead, and effective optimization. To overcome these limitations, the authors propose DVM, a just-in-time compiler that introduces a novel bytecode-based execution mechanism directly targeting NPUs. DVM employs a bytecode virtual machine to efficiently compile dynamic operator instances and integrates symbolic shape inference from static graphs with runtime fusion strategies from dynamic graphs, enabling on-the-fly operator compilation and multi-granularity fusion. Experimental results demonstrate that DVM achieves up to an 11.77× speedup in operator and model execution compared to TorchInductor, PyTorch-eager, and MindSpore-graph-O0, while reducing compilation time by up to five orders of magnitude.
This study addresses the functional mismatches arising from the lack of hardware-level verification in compilation mappings for machine learning accelerators. To this end, it proposes BOLT, a formal verification framework that aligns loop structures with associated data layouts. Notably, BOLT introduces the first coarse-grained intrinsic verification method that operates without requiring additional auxiliary information. Furthermore, by leveraging synchronization skeleton and layout sketch template techniques, the framework achieves end-to-end correctness proofs for compiler-to-accelerator mappings. The prototype system successfully verifies the correctness of complex compilation mappings on two open-source ML accelerators, thereby providing a reliable foundation for hardware-software co-design.
本文针对机器学习编译器中张量内存布局选择问题,提出了一种基于数据流图的组合优化方法,并设计了多项式时间算法和加权MaxSAT编码来解决该问题。
This work addresses the challenge of efficiently generating vector-length-agnostic (VLA) machine learning code for scalable vector instruction sets such as Arm SVE, where unknown vector lengths at compile time hinder traditional compilers. The authors present the first end-to-end VLA support in MLIR/IREE, introducing a vector-length-aware compact data layout and unifying dynamic tiling, operator fusion, and scalable vectorization within a single compilation framework. Evaluated on Arm CPUs, the generated SVE code achieves up to 1.45× speedup over IREE’s NEON implementation, outperforms multiple frameworks in the PyTorch ecosystem, and demonstrates strong scalability with increasing vector lengths in simulation, effectively balancing performance and hardware portability.
本文提出了一种符号程序分析方法,用于机器学习内核的基本块计数,相比动态插桩技术,显著减少了运行时间。
This study addresses the high maintenance costs of compiler backends and the difficulty of adapting to emerging hardware by proposing the AI Lowering paradigm. This approach leverages large language model (LLM) agents to directly compile Triton kernels into PTX code, thereby bypassing conventional optimization pipelines. Furthermore, it substantially extends the Volta validator to accommodate Blackwell architecture features. Experimental results demonstrate that the proposed method achieves performance improvements ranging from 0.83× to 3.34× across diverse GPU platforms, significantly reducing the software deployment overhead associated with new chip architectures.
This work addresses the challenge that lightweight compilers and source-to-source tools struggle to reuse the sophisticated inlining heuristics of mature compilers like GCC or LLVM due to their reliance on complex intermediate representations and analysis infrastructures. To bridge this gap, the paper introduces the first portable inlining prediction framework that leverages diagnostic outputs from production compilers as supervision signals. By extracting call-site features through AST normalization and constructing a lightweight structured IR, the approach trains tabular models—such as CatBoost—that can be directly compiled into pure C code without runtime dependencies. Evaluated on a dataset of 330,000 call sites, the model achieves a ROC-AUC of 0.928 and PR-AUC of 0.713; with threshold tuning, it attains an F1 score of 0.729 while reducing the false positive rate to 0.084, thereby enabling the first practical transfer of industrial-grade inlining decisions to resource-constrained systems.
This study addresses the challenges of backend inference and deployment reliability arising from the separation of compilation and execution in machine learning stacks. To this end, it proposes the first Rust-based unified architecture that deeply integrates the compiler and runtime. Methodologically, a three-level intermediate representation (IR) is designed to enable transparent multi-backend scheduling and distributed scaling, supporting 14 device types and customized code generation paths alongside strict fault-handling mechanisms. The system further incorporates TCP/RDMA communication, mixed-precision training, and sparse linear algebra techniques to enhance overall capabilities. Experimental results demonstrate that the proposed system maintains 100% accuracy on MiniLM and MNIST benchmarks while surpassing mainstream frameworks such as PyTorch in throughput and significantly reducing latency.
This study addresses the performance limitations of GPU kernels generated by existing compilers and the shortcomings of LLM-based optimization approaches, which often overlook model structure and lack end-to-end verification. We propose a multi-agent collaborative framework that treats compiled models as structured artifacts, focusing specifically on Triton sub-kernel optimization. This work introduces a novel schedule-aware agent search mechanism, integrated with vendor library call protection and a four-stage gated cascaded verification pipeline encompassing static analysis, correctness checking, and performance gating to ensure both safety and efficacy. Evaluated on the KernelBench benchmark, our method achieves average speedups of 1.40×, 1.15×, and 1.07× at the L1, L2, and L3 levels, respectively, compared to torch.compile, significantly enhancing inference efficiency.