Score
Designs, builds, and evaluates compiler internals and end-to-end compilation pipelines that perform operation fusion and related transformations to improve locality and performance. This includes designing compiler intermediate representations and analyses, implementing compiler transformations, and creating heterogeneous and connectivity‑aware compilation strategies that reason about data movement, device placement, and interconnect when generating code for multiple targets.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.
This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.
This work addresses the performance limitations of traditional pointer analysis by proposing a decoupled acceleration paradigm. Instead of tightly coupling simplification rules with the analysis itself—a common drawback of existing offline approaches—the method applies general-purpose, semantics-preserving compiler optimizations to the intermediate representation (IR) prior to analysis. This modular, analysis-agnostic strategy enhances efficiency without requiring modifications to the pointer analysis algorithm, thereby supporting seamless integration with diverse analyses. Empirical evaluation across multiple benchmark programs and three mainstream pointer analyses demonstrates speedups of up to 3.14× and memory reductions of up to 1.94×, all while largely preserving precision.
Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.
This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.
This work proposes a novel approach within the BuildIt system to circumvent the engineering complexity and high cost of traditional program optimization, which relies on backward data-flow analysis to obtain future execution information. Instead of performing backward analysis, the method introduces oracle variables that enable multiple forward executions to predict future program behavior. This is the first application of oracle variables to support optimizations requiring future information, eliminating the need for intermediate representations or custom parsers. Consequently, it significantly reduces implementation overhead for embedded domain-specific languages (DSLs) built atop C++. Integrated with staged compilation and code generation targeting C, C++, and CUDA, experimental results demonstrate that the approach not only enhances performance but also substantially decreases the engineering effort required for DSL development.
CPU-GPU data transfers in AI-based code generation impose severe compilation and iterative latency bottlenecks. Method: This paper introduces the first theoretical framework for GPU-native compilation, systematically proposing and analyzing three paradigms—parallel traditional compilation, neural compilation, and hybrid compilation. Contributions/Results: (1) A probabilistic formal verification mechanism that jointly optimizes accuracy and parallelism; (2) Theoretical characterization of latency, energy, and correctness trade-offs across all three paradigms, along with provable co-optimization pathways; (3) A deployable hybrid architecture ensuring strict correctness. Experimental and theoretical analysis demonstrates up to 10–100× speedup in end-to-end code iteration latency: traditional GPU compilation accelerates by 2–5×, neural compilation by 10–100×, and the hybrid approach achieves practical efficiency without compromising formal correctness.
This work addresses the challenge that lightweight compilers and source-to-source tools struggle to reuse the sophisticated inlining heuristics of mature compilers like GCC or LLVM due to their reliance on complex intermediate representations and analysis infrastructures. To bridge this gap, the paper introduces the first portable inlining prediction framework that leverages diagnostic outputs from production compilers as supervision signals. By extracting call-site features through AST normalization and constructing a lightweight structured IR, the approach trains tabular models—such as CatBoost—that can be directly compiled into pure C code without runtime dependencies. Evaluated on a dataset of 330,000 call sites, the model achieves a ROC-AUC of 0.928 and PR-AUC of 0.713; with threshold tuning, it attains an F1 score of 0.729 while reducing the false positive rate to 0.084, thereby enabling the first practical transfer of industrial-grade inlining decisions to resource-constrained systems.