Score
Designs, implements, and evaluates compiler passes, lowering and code‑generation pipelines targeting hardware accelerators, including XLA/MLIR-based optimizations such as primitive fusion, buffer layout and allocation, operator lowering, loop scheduling, and backend code emission. Builds and tunes integration and autotuning components, and analyzes correctness and performance impacts of XLA compiler internals, compilation parameters, and compiler optimization passes.
Automated optimization of complex nested loops on modern hardware remains challenging due to the combinatorial complexity of legal and profitable loop transformations. Method: This paper proposes an LLM-guided closed-loop compilation optimization framework that leverages a general-purpose large language model—without fine-tuning or in-context examples—as an intelligent agent. Grounded in the polyhedral model, the LLM generates loop transformation schedules, which are iteratively refined using real-time compiler feedback on both performance speedup and semantic correctness. Contribution/Results: To our knowledge, this is the first zero-shot, feedback-driven autonomous scheduling approach, introducing the embodied intelligence paradigm into compiler optimization. Evaluated on the PolyBench benchmark, it achieves an average 2.66× speedup per run and up to 3.54× after five iterations—substantially outperforming state-of-the-art tools such as Pluto—demonstrating the feasibility and superiority of LLM–compiler co-optimization for efficient, reliable automatic loop optimization.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
Existing automated pass-tuning methods assume linear pass sequences, rendering them incompatible with LLVM’s new Pass Manager—which employs a nested hierarchical structure—thus yielding syntactically invalid pipelines. Method: We propose the first automated tuning framework for nested optimization pipelines: (i) we formally define a context-free grammar specifying syntactically valid pipelines; (ii) we introduce a forest-based data structure to natively represent nested pass topologies; and (iii) we design a structure-aware, co-evolutionary genetic algorithm augmented with a local refinement mechanism. Results: Evaluated across seven benchmark suites, our generated optimization policies reduce instruction count by 13.62% on average over `opt -Oz`, demonstrating substantial improvement in discovering high-performance, syntactically valid pass combinations under complex structural constraints.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
Manual pragma configuration in high-level synthesis (HLS) suffers from low efficiency and an exponentially large search space. Method: This paper proposes the first nonlinear programming (NLP)-based automated pragma insertion framework, jointly optimizing loop-level pragmas—including pipelining, function unit replication, and data caching. It innovatively models discrete pragma configurations as continuous, differentiable variables and constructs analytical performance/resource models with theoretical lower-bound guarantees, solved globally via NLP. Integrated with pragma semantic analysis and the Merlin compiler, and augmented by design-space pruning, the framework explores billion-scale configurations within seconds to minutes. Contribution/Results: Experimental evaluation shows kernel performance approaching hand-tuned implementations, resource estimation error <8%, and latency lower-bound error ≤12%.
This study addresses the challenges of pass ordering, selection, and performance analysis in LLVM’s -O3 optimization pipeline by systematically decomposing the optimization process into per-pass prefixes. Conducting 84,750 noise-controlled experiments across 30 PolyBench/C kernels, the work comprehensively evaluates multidimensional metrics including execution time, compilation overhead, binary size, hardware counters, and energy consumption. It reveals for the first time that optimization benefits are highly concentrated in a few critical passes, that -O3 is rarely Pareto-optimal across programs, and that IR instruction count poorly predicts runtime performance. The study establishes a theoretical upper bound on phase-ordering interference loss and demonstrates that 6.6–9.7% of passes degrade performance, 84.8% of the full pipeline is needed to achieve 80% of peak speedup, and runtime-synchronized optimizations can reduce energy usage by 30–60%, providing an empirical foundation for autotuning and cost-model calibration.
This work addresses the subtle microarchitectural performance inefficiencies often introduced by modern compiler optimizations, which can lead to significant yet overlooked performance losses. The authors propose a top-down differential analysis methodology that systematically identifies and categorizes the root causes of such optimization defects by integrating fine-grained microarchitectural performance counter sampling with cross-compiler (GCC/Clang) binary comparisons. Innovatively combining top-down microarchitectural analysis with differential testing, the approach further introduces a portable binary patching framework to precisely locate and rectify inefficient code segments. Empirical evaluation demonstrates that the method effectively uncovers substantial but commonly neglected performance discrepancies between GCC and Clang and successfully recovers performance through targeted binary patches.
This work proposes a hybrid concrete-symbolic interpretation method to efficiently verify semantic equivalence between original and optimized programs in MLIR, ensuring the correctness of optimization transformations. The approach supports diverse syntactic, scheduling, and memory representations and theoretically achieves linear-time complexity for equivalence checking. Building upon this method, the authors develop a formal verifier for a subset of MLIR and successfully apply it to the AMD MLIR-AIR and MLIR-AIE toolchains as well as the standard mlir-opt infrastructure. Evaluation across hundreds of benchmark variants demonstrates the verifier’s effectiveness in validating optimization pipelines, significantly enhancing the reliability of compiler optimizations within the MLIR ecosystem.
This work addresses the challenge of efficiently exploiting parallelism and mitigating latency from hierarchical memory and explicit data movement in edge AI kernels deployed on resource-constrained devices. Building upon the MLIR compilation framework and leveraging kernels generated by Triton/Inductor, the study systematically evaluates three compiler-level optimizations: vectorization (Vec), hardware context-level multithreading (MT), and ping-pong double buffering (DB). Through a novel ablation staircase methodology, the authors uniquely isolate and quantify the individual and synergistic performance contributions of these techniques under varying compute-to-memory intensity regimes. The findings reveal that vectorization primarily accelerates bandwidth-bound kernels, multithreading yields significant gains after amortizing scheduling overhead, and double buffering provides additional speedup when computation and data transfer can be effectively overlapped.