Score
Designs, implements, and evaluates compiler and code-generation systems and optimization passes that transform high-level machine-learning representations into efficient low-level code. This includes building code-generation pipelines and tooling (e.g., MLIR/LLVM), kernel and kernel-fusion generators, code-transformation and optimization pipelines, automated and AI-assisted (including neural) code-generation models and workflows, and the integration and extensibility of code-generation systems.
Traditional compilers face limitations in development accessibility, optimization capabilities, and application scope. This work proposes the first multidimensional classification framework for large language model (LLM)-driven compilation, offering a systematic survey of existing research through four analytical dimensions: design philosophy, methodology, level of code abstraction, and task type. The study identifies three core design paradigms—Selector, Translator, and Generator—and highlights three transformative directions: democratizing compiler development, discovering novel optimization strategies, and expanding functional boundaries. It further argues that hybrid systems represent a critical pathway forward and provides a technical roadmap for building correct, scalable, and intelligent compilation tools.
This work proposes a novel approach to program analysis and optimization leveraging large language models (LLMs). Addressing the challenge of effectively integrating source code and intermediate representation (IR) information—a limitation in existing methods—it introduces LLMCompiler, pre-trained on IR, and employs a chunked embedding and aggregation strategy to produce unified program-level embeddings. By innovatively unifying the semantics of source code and IR, the method achieves a 1.54% error rate on algorithm classification, representing a 12% improvement over the current state of the art. It also attains competitive accuracy in heterogeneous device mapping, significantly advancing the application of LLMs in program understanding and optimization.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
Existing black-box neural-network-based AI code generation methods lack formal guarantees for the correctness and legality of loop scheduling transformations. Method: We propose the first end-to-end, polyhedral-driven framework supporting formal correctness guarantees, tightly integrating machine learning optimization with compilation theory. Our approach constructs a verifiable program transformation space, introduces differentiable modeling of legality constraints, and jointly generates training data annotated with machine-checkable proofs. Contribution/Results: All generated scheduling transformations are formally proven to preserve semantic equivalence under the polyhedral model. The framework incurs negligible runtime overhead and demonstrates strong generalization across diverse hardware platforms—including CPUs, GPUs, and FPGAs—as well as canonical loop optimization scenarios. An open-source implementation is publicly available.
This study investigates the feasibility and limitations of AI-driven code optimization. We systematically evaluate three traditional compilers (GCC, LLVM, CETUS) against two large language models (LLMs)—CodeLlama-70B and DeepSeek-Coder—across performance and functional correctness. To this end, we introduce the first LLM-specific benchmark and automated verification framework for compiler optimizations, enabling multi-dimensional joint assessment of speedup and correctness. We further propose compilation-strategy-embedded prompting techniques—DIP (Domain-Informed Prompting) and CoT (Chain-of-Thought)—to enhance LLM reasoning for optimization tasks. Experimental results show that CodeLlama-70B achieves a 1.75× speedup on average, surpassing the best-performing compiler CETUS (1.67×); however, it exhibits high error rates on large-scale code, underscoring the necessity of rigorous verification. Our core contributions include: (1) a reproducible evaluation methodology for LLM-based compilation, (2) empirical evidence of prompt engineering’s critical role in LLM-driven optimization, and (3) the establishment of an “LLM + formal verification” co-optimization paradigm.
This work systematically evaluates the capability boundaries of Llama 3.1 405B on natural language-to-multilingual executable code generation, focusing on algorithmic problem solving and fundamental data structure tasks. To address limitations in cross-lingual code synthesis and robustness, we propose a synergistic approach integrating context-aware prompt engineering, multilingual code fine-tuning, and dynamic test-time validation. On standard benchmarks—including HumanEval and MBPP—our method achieves state-of-the-art performance, attaining >82% pass@1 accuracy on medium-difficulty algorithmic problems. We further provide the first empirical evidence that Llama 3.1 405B exhibits strong generalization on classical computer science problems (e.g., sorting, graph traversal), yet its accuracy drops substantially in frontier domains such as quantum computing and bioinformatics. These findings offer rigorous empirical support and a concrete technical pathway for large language model–driven programming assistance and computational education.
This work addresses the interoperability challenge between GCC and LLVM compiler intermediate representations (IRs), which stems from their semantic and structural differences. To bridge this gap, the authors propose IRIS-14B, the first large language model specifically designed for IR-to-IR translation. Built upon a 14-billion-parameter Transformer architecture, IRIS-14B leverages supervised fine-tuning to learn the mapping between GIMPLE and LLVM IR derived from the same C source code, enabling high-fidelity automatic translation. Experimental results demonstrate that IRIS-14B substantially outperforms existing open-source large models on real-world C programs and competitive programming tasks, achieving up to a 44-percentage-point improvement in accuracy. This study provides the first empirical validation of large language models as effective and feasible interoperability layers within neuro-symbolic hybrid compilation frameworks.
This paper investigates the trade-off between search efficiency and optimization effectiveness in automated code optimization, specifically addressing whether fixed-order application of code transformations can substantially reduce the search space without significantly compromising performance gains. To this end, we propose a data-driven empirical methodology: generating random programs, executing randomized optimization sequences, measuring runtime performance, and constructing a large-scale experimental dataset to statistically characterize inter-transform interactions. Our results demonstrate that several high-performance fixed optimization sequences exist—achieving over 95% of the optimal performance gain while reducing the search space by more than 90% on average. This finding provides empirically validated theoretical support and practical design principles for compiler auto-optimization strategies.
This work addresses the limited understanding of large language models’ (LLMs’) capability in generating efficient code for CPU-oriented high-performance computing and domestic heterogeneous architectures. The authors present CodegenBench, the first systematic benchmark evaluating mainstream LLMs on their ability to generate performant parallel code across three hardware platforms: x86_64, Sunway, and Kunpeng, covering 106 BLAS routines and 60 computational kernels. Through automated cross-architecture compilation, runtime profiling, and performance analysis, the study reveals that while LLMs can produce highly efficient code on well-documented x86_64 systems, their performance degrades significantly on specialized architectures with sparse documentation. The models achieve best results on tasks of moderate complexity and short implementation length. The released dataset and evaluation framework fill a critical gap in LLM research for domestic supercomputing platforms.
Automated optimization of complex nested loops on modern hardware remains challenging due to the combinatorial complexity of legal and profitable loop transformations. Method: This paper proposes an LLM-guided closed-loop compilation optimization framework that leverages a general-purpose large language model—without fine-tuning or in-context examples—as an intelligent agent. Grounded in the polyhedral model, the LLM generates loop transformation schedules, which are iteratively refined using real-time compiler feedback on both performance speedup and semantic correctness. Contribution/Results: To our knowledge, this is the first zero-shot, feedback-driven autonomous scheduling approach, introducing the embodied intelligence paradigm into compiler optimization. Evaluated on the PolyBench benchmark, it achieves an average 2.66× speedup per run and up to 3.54× after five iterations—substantially outperforming state-of-the-art tools such as Pluto—demonstrating the feasibility and superiority of LLM–compiler co-optimization for efficient, reliable automatic loop optimization.
Large language models (LLMs) excel at code generation, yet the impact of their compressed variants—such as those produced via quantization or knowledge distillation—on token representations of programming languages remains poorly understood, hindering deployment quality. This work systematically investigates how LLM tokenizers encode programming languages and introduces a novel “cold-start probability” analysis method that operates without explicit prompting. By integrating lexical distribution analysis, keyword coverage, and multidimensional evaluation metrics, the study provides the first comprehensive characterization of the subtle effects of compression strategies—including quantization, knowledge distillation, model scaling, and task-specific fine-tuning—on code token representations. The findings offer both theoretical grounding and empirical guidance for deploying high-quality, efficient code generation models.