Score
Designs, implements, and tunes compiler backend components and pipelines, including IR lowering passes, backend-specific optimization algorithms, and register allocation techniques. Builds and analyzes compiler toolchain elements such as cost and compute allocation models, feedback loops, and performance analyses to drive backend optimization, pipeline design, and overall compiler backend engineering.
This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.
This work proposes a novel approach within the BuildIt system to circumvent the engineering complexity and high cost of traditional program optimization, which relies on backward data-flow analysis to obtain future execution information. Instead of performing backward analysis, the method introduces oracle variables that enable multiple forward executions to predict future program behavior. This is the first application of oracle variables to support optimizations requiring future information, eliminating the need for intermediate representations or custom parsers. Consequently, it significantly reduces implementation overhead for embedded domain-specific languages (DSLs) built atop C++. Integrated with staged compilation and code generation targeting C, C++, and CUDA, experimental results demonstrate that the approach not only enhances performance but also substantially decreases the engineering effort required for DSL development.
Compiler optimization reports are highly technical, difficult to interpret, and challenging to operationalize. To address this, we propose CompilerGPT—the first end-to-end framework that deeply integrates large language models (LLMs) into the compiler optimization feedback loop. CompilerGPT synergistically combines GPT-4o and Claude Sonnet with static analysis, structured prompt engineering, test-driven feedback, and multi-round iterative execution to enable cross-compiler (Clang/GCC) report parsing and executable code rewriting. Its core innovation lies in establishing a verifiable, automated optimization workflow that closes the loop from report comprehension to measurable performance improvement. Evaluated on five benchmark programs, CompilerGPT achieves up to 6.5× runtime speedup, empirically demonstrating the feasibility and effectiveness of LLM-driven automation for compiler optimizations.
本文针对编译器优化启发式算法导致的性能差和编译时间不可预测问题,提出验证编译阶段的性能与编译时间属性的方法,并以内联扩展为例进行了证明。
Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.
This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.
This paper investigates the trade-off between search efficiency and optimization effectiveness in automated code optimization, specifically addressing whether fixed-order application of code transformations can substantially reduce the search space without significantly compromising performance gains. To this end, we propose a data-driven empirical methodology: generating random programs, executing randomized optimization sequences, measuring runtime performance, and constructing a large-scale experimental dataset to statistically characterize inter-transform interactions. Our results demonstrate that several high-performance fixed optimization sequences exist—achieving over 95% of the optimal performance gain while reducing the search space by more than 90% on average. This finding provides empirically validated theoretical support and practical design principles for compiler auto-optimization strategies.
This work addresses the limitation of conventional hardware compilation flows, wherein pipeline optimization is deferred to the backend, resulting in the loss of high-level structural information and suboptimal global optimization. To overcome this, the paper introduces, for the first time, an explicit pipeline-aware compiler pass at the intermediate representation (IR) level. By modeling legality constraints for register relocation, integrating learning-driven timing prediction, and formulating the timing-constrained relocation problem as a global minimum-cost flow problem, the approach enables timing-aware RTL generation. Implemented within the CIRCT framework, the method significantly reduces critical-path delay, power, and area across both open-source and commercial designs, while also providing backend retiming with a superior initial structure.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
Accelerator Design Language (ADL) compilers suffer from unpredictable hardware performance due to the semantic gap between high-level abstractions and low-level implementations, compounded by reliance on heuristic optimizations. This work introduces Petal, the first tool enabling cycle-accurate, interpretable performance analysis for Calyx-based accelerators. Petal bridges the abstraction gap via three key techniques: source-code instrumentation, RTL-simulation trace collection, and a novel reverse-mapping algorithm that associates low-level timing events in synthesized hardware with high-level control-flow constructs. Crucially, it abandons the “compiler-perfection” assumption, instead exposing how concrete compilation decisions impact end-to-end latency. Evaluated on multiple real-world accelerator designs, Petal identifies subtle, manually elusive bottlenecks—enabling targeted manual optimizations that reduce total execution cycles by up to 46.9% for one application.