model compiler implementation

Implements compilers for machine-learning models by designing and building the pipeline that lowers high-level model graphs into optimized intermediate representations and low-level code, including passes for operator fusion, scheduling, memory/layout transforms, and other optimizations. Produces target-specific code generators and runtime integration (XLA-/TVM-style) to emit efficient kernels and binaries for CPUs, GPUs, and accelerators.

modelcompilerimplementation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization

Jul 07, 2025
SA
Samira Ahmadifarsani
🏛️ Technical University of Munich | TU Wien

Integrating custom hardware accelerators—particularly GEMM-based ones—into mainstream ML compilers remains challenging due to tight coupling between accelerator-specific optimizations and compiler internals. Method: This paper proposes a high-level, low-intrusion integration methodology for TVM that abstracts hardware scheduling interfaces to decouple accelerator characteristics from compiler implementation. Leveraging the CoSA design-space exploration framework, it automates hardware-aware scheduling optimizations—including tensor tiling, non-uniform mapping, and double buffering—without modifying TVM’s core infrastructure. Contribution/Results: Evaluated on the Gemmini accelerator, the approach achieves performance on par with hand-optimized toolchains while significantly improving developer productivity and cross-model/cross-architecture portability. The abstraction enables seamless reuse of scheduling policies across diverse accelerator microarchitectures and neural network workloads, reducing integration effort from weeks to hours.

Automating efficient tensor scheduling for deep learning accelerators is challengingExisting frameworks require deep compiler knowledge and offer partial solutionsIntegrating custom accelerators into ML compilers is complex

This work addresses the challenges posed by dynamic AI models—whose varying tensor shapes and control flows hinder existing compilers from simultaneously achieving fast compilation, low memory overhead, and effective optimization. To overcome these limitations, the authors propose DVM, a just-in-time compiler that introduces a novel bytecode-based execution mechanism directly targeting NPUs. DVM employs a bytecode virtual machine to efficiently compile dynamic operator instances and integrates symbolic shape inference from static graphs with runtime fusion strategies from dynamic graphs, enabling on-the-fly operator compilation and multi-granularity fusion. Experimental results demonstrate that DVM achieves up to an 11.77× speedup in operator and model execution compared to TorchInductor, PyTorch-eager, and MindSpore-graph-O0, while reducing compilation time by up to five orders of magnitude.

compilation overheaddynamic AI modelsoperator fusion

This study addresses the functional mismatches arising from the lack of hardware-level verification in compilation mappings for machine learning accelerators. To this end, it proposes BOLT, a formal verification framework that aligns loop structures with associated data layouts. Notably, BOLT introduces the first coarse-grained intrinsic verification method that operates without requiring additional auxiliary information. Furthermore, by leveraging synchronization skeleton and layout sketch template techniques, the framework achieves end-to-end correctness proofs for compiler-to-accelerator mappings. The prototype system successfully verifies the correctness of complex compilation mappings on two open-source ML accelerators, thereby providing a reliable foundation for hardware-software co-design.

Compiler-to-Accelerator MappingsFormal VerificationFunctional Equivalence

This work addresses the challenge of efficiently generating vector-length-agnostic (VLA) machine learning code for scalable vector instruction sets such as Arm SVE, where unknown vector lengths at compile time hinder traditional compilers. The authors present the first end-to-end VLA support in MLIR/IREE, introducing a vector-length-aware compact data layout and unifying dynamic tiling, operator fusion, and scalable vectorization within a single compilation framework. Evaluated on Arm CPUs, the generated SVE code achieves up to 1.45× speedup over IREE’s NEON implementation, outperforms multiple frameworks in the PyTorch ecosystem, and demonstrates strong scalability with increasing vector lengths in simulation, effectively balancing performance and hardware portability.

compiler code generationdata layoutML compilation

Latest Papers

What's happening recently
View more

This study addresses the high maintenance costs of compiler backends and the difficulty of adapting to emerging hardware by proposing the AI Lowering paradigm. This approach leverages large language model (LLM) agents to directly compile Triton kernels into PTX code, thereby bypassing conventional optimization pipelines. Furthermore, it substantially extends the Volta validator to accommodate Blackwell architecture features. Experimental results demonstrate that the proposed method achieves performance improvements ranging from 0.83× to 3.34× across diverse GPU platforms, significantly reducing the software deployment overhead associated with new chip architectures.

AI compilercompiler backendlarge language models

This work addresses the challenge that lightweight compilers and source-to-source tools struggle to reuse the sophisticated inlining heuristics of mature compilers like GCC or LLVM due to their reliance on complex intermediate representations and analysis infrastructures. To bridge this gap, the paper introduces the first portable inlining prediction framework that leverages diagnostic outputs from production compilers as supervision signals. By extracting call-site features through AST normalization and constructing a lightweight structured IR, the approach trains tabular models—such as CatBoost—that can be directly compiled into pure C code without runtime dependencies. Evaluated on a dataset of 330,000 call sites, the model achieves a ROC-AUC of 0.928 and PR-AUC of 0.713; with threshold tuning, it attains an F1 score of 0.729 while reducing the false positive rate to 0.084, thereby enabling the first practical transfer of industrial-grade inlining decisions to resource-constrained systems.

compiler optimizationfunction inliningheuristics transfer

This study addresses the challenges of backend inference and deployment reliability arising from the separation of compilation and execution in machine learning stacks. To this end, it proposes the first Rust-based unified architecture that deeply integrates the compiler and runtime. Methodologically, a three-level intermediate representation (IR) is designed to enable transparent multi-backend scheduling and distributed scaling, supporting 14 device types and customized code generation paths alongside strict fault-handling mechanisms. The system further incorporates TCP/RDMA communication, mixed-precision training, and sparse linear algebra techniques to enhance overall capabilities. Experimental results demonstrate that the proposed system maintains 100% accuracy on MiniLM and MNIST benchmarks while surpassing mainstream frameworks such as PyTorch in throughput and significantly reducing latency.

Distributed RuntimeGraph CompilationMachine Learning Infrastructure

This study addresses the performance limitations of GPU kernels generated by existing compilers and the shortcomings of LLM-based optimization approaches, which often overlook model structure and lack end-to-end verification. We propose a multi-agent collaborative framework that treats compiled models as structured artifacts, focusing specifically on Triton sub-kernel optimization. This work introduces a novel schedule-aware agent search mechanism, integrated with vendor library call protection and a four-stage gated cascaded verification pipeline encompassing static analysis, correctness checking, and performance gating to ensure both safety and efficacy. Evaluated on the KernelBench benchmark, our method achieves average speedups of 1.40×, 1.15×, and 1.07× at the L1, L2, and L3 levels, respectively, compared to torch.compile, significantly enhancing inference efficiency.

deep learning compilerend-to-end verificationGPU kernel optimization

Hot Scholars

ZZ

Ziqi Zhu

PhD student, University of Science and Technology of China
AI for ScienceComputer Vision
OB

Oliver Bringmann

Professor of Embedded Systems, Eberhard Karls Universität Tübingen, Germany
Embedded System DesignSystem Modeling and SimulationAutomotive ElectronicsSafety-critical Systems
US

Ulf Schlichtmann

Professor for Electronic Design Automation, Technical University of Munich
Electronic Design AutomationReliability/Robustness/ResilienceEmbedded SystemsMicrofluidic Biochips
MZ

Minfan Zhao

University of Science and Technology of China
computer vision