build compiler backends

Designs and implements compiler backends by defining intermediate representations and transformation passes, implementing lowering, code generation and scheduling, and translating DSLs into executable IR such as XLA HLO. Integrates domain-specific optimizations, generates backend-specific code, applies cross-backend optimizations, and ensures correct and portable execution semantics when mapping workloads to target hardware.

buildcompilerbackends

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$227K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

TPDE: A Fast Adaptable Compiler Back-End Framework

May 28, 2025
TS
Tobias Schwarz
🏛️ Technical University of Munich

Existing compilers (e.g., LLVM) prioritize optimization over low-latency compilation, incurring substantial IR transformation overhead; custom backends suffer from high development costs and poor cross-architecture portability. This paper proposes a lightweight, single-pass, adaptive JIT backend framework that directly consumes source IR in SSA form and unifies instruction selection, register allocation, and machine code emission. It introduces the first IR-agnostic adapter mechanism, enabling zero-transformation integration of heterogeneous IRs—including LLVM IR, WebAssembly bytecode, and database query plans. Semantic-driven, architecture-independent optimizations are performed natively, with built-in support for x86-64 and AArch64. Evaluated on SPECint 2017, our compiler achieves 8–24× faster compilation than LLVM at -O0, while matching its runtime performance. In WebAssembly and database JIT scenarios, end-to-end compilation latency is significantly reduced.

Eliminating extra IR translation steps in compiler frameworksFast machine code generation for low-latency JIT compilationSimplifying multi-architecture support in custom code generators

This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.

Compiler Intermediate RepresentationCross-toolchain InteroperabilityGIMPLE

A Unified Framework for Automated Code Transformation and Pragma Insertion

May 05, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.

AutomationCode ModificationSimplification

Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.

compiler correctnesscompiler optimizationhybrid systems

A shared compilation stack for distributed-memory parallelism in stencil DSLs

Apr 02, 2024
GB
George Bisbas
🏛️ Imperial College London | The University of Edinburgh | Technische Universität Berlin | University of Cambridge

High-performance computing (HPC) stencil domain-specific language (DSL) compilers suffer from high development costs, poor infrastructure reuse, and low maintainability due to isolated, ad hoc designs. To address these challenges, this paper proposes MLIR-HPC, a dedicated extensible compiler framework for HPC built on the MLIR infrastructure. Our method introduces three key innovations: (1) a novel message-passing abstraction for distributed-memory systems that uniformly models communication semantics; (2) a distributed stencil intermediate representation (IR) supporting automated communication generation and cross-DSL optimization passes; and (3) seamless integration with three major DSL backends—Devito, PSyclone, and Open Earth Compiler—enabling shared compilation stack infrastructure. Evaluated across heterogeneous supercomputing architectures, the framework supports all three stencil DSLs using a unified core, achieving industrial-grade compilation efficiency and execution performance. Results demonstrate significantly enhanced sustainability, reusability, and evolutionary capability for HPC DSL compilers.

Develop shared compiler infrastructure for distributed-memory parallelism in stencil DSLs.Enable high-performance executables across multiple HPC stencil-DSL compilers.Reduce development and maintenance costs of DSL compilers in HPC.

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified framework in existing domain-specific language (DSL) compilers, which leads to redundant development, maintenance challenges, and difficulty meeting production-grade requirements. The paper presents the first fully MLIR-based NumPy-like DSL, featuring native implementation of both front-end parsing and semantic analysis within MLIR. It introduces a novel dialect-agnostic type checker and a parallelism-first lowering strategy that seamlessly integrates with MLIR’s dataflow dialects. By doing so, this approach not only advances the standardization of DSLs within the MLIR ecosystem but also demonstrates strong performance on real-world Fortran applications in domains such as weather modeling and computational fluid dynamics.

code reusecompiler frameworksDSL compilers

Traditional compilers face limitations in applying equality saturation–based optimizations due to their reliance on a single abstraction level and their inability to preserve discovered equivalences across subsequent compiler transformations, leading to phase-ordering problems. This work proposes natively embedding e-graphs into the compiler’s intermediate representation (IR), introducing eqsat as a first-class dialect within MLIR. This design enables the persistent maintenance of equality saturation state throughout the compilation pipeline and supports interleaved execution with other transformations. Implemented within the xDSL framework, the approach unifies equality saturation optimization across multiple IR abstraction levels, significantly enhancing both the power and flexibility of compiler optimizations.

compiler optimizatione-graphequality saturation

This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.

abstractionscode performance optimizationhigh-performance computing

This work addresses the challenge of programming hard intellectual property (hard IP) blocks—such as tensor slices—in domain-specific FPGAs, which are typically inaccessible to high-level synthesis (HLS) and require inefficient manual RTL integration. The authors propose a compiler-agnostic approach that leverages HLS black-box mechanisms at the architectural level: hard IP modules are encapsulated as RTL black boxes and modeled as C-level schedulable operators with explicit latency and initiation interval constraints. This enables standard HLS tools, such as AMD Vitis HLS, to directly invoke hard IPs from C/C++ code without compiler modifications or handcrafted co-design. Experiments on a tensor-slice FPGA demonstrate that the proposed method yields designs with lower area-delay product compared to behavioral HLS baselines, while achieving hardware efficiency comparable to hand-written RTL but with substantially improved development productivity.

Domain-specific architectureFPGA hardblocksHigh-Level Synthesis

Manually configuring linters requires expert knowledge and struggles to adapt across multiple programming languages, coding standards, and tooling ecosystems, leading to high maintenance overhead. This work proposes LintCFG, the first approach to apply compiler design principles to automated linter configuration generation. It introduces a tool-agnostic domain-specific language (DSL) to structurally encode coding rules and leverages large language models to automatically compile natural language specifications into concrete linter configurations, enabling end-to-end automation across languages, standards, and tools. Evaluated on Java Checkstyle tasks, the DSL achieves over 90% precision and recall in rule representation, with fine-grained configuration generation exceeding 70% accuracy—more than doubling the performance of baseline methods. User studies confirm significant gains in developer productivity, and the approach successfully generalizes to JavaScript ESLint scenarios.

coding standardsconfiguration maintenancelinter configuration

Hot Scholars

TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
PS

Philipp Schaad

ETH Zurich
HPCcompilersprogram visualizationparallel programming
JC

Jeronimo Castrillon

Professor, TU Dresden
CompilersHeterogeneous SystemsEmerging ComputingDomain-specific Languages