Score
Design and implement physical memory representations and transformation pipelines that map logical data structures into concrete in-memory layouts and generate alternative layouts. Build analyses and compiler or toolchain transformations that optimize those layouts to minimize copying and packing, improve alignment and efficient kernel access, and produce architecture-alapted layout variants during compilation.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
To address the optimization limitations imposed by tight coupling between data layout and computation on GPUs, this paper proposes a layout-agnostic computational abstraction paradigm: computations are first expressed in a layout-decoupled form, and hierarchical parallel index expressions are then automatically derived via layout specifications. The core contribution is the first end-to-end “layout → index expression” mapping mechanism, enabling layout-driven code generation and cross-compiler optimization exploration. We design a custom layout specification language and integrate it with MLIR, Triton, and CUDA templates to build an index derivation engine. Experimental evaluation demonstrates that the generated code achieves performance on par with hand-optimized Triton kernels. Furthermore, the approach is validated for generality and efficiency across both MLIR-based and CUDA-based compilation ecosystems.
Manual pragma configuration in high-level synthesis (HLS) suffers from low efficiency and an exponentially large search space. Method: This paper proposes the first nonlinear programming (NLP)-based automated pragma insertion framework, jointly optimizing loop-level pragmas—including pipelining, function unit replication, and data caching. It innovatively models discrete pragma configurations as continuous, differentiable variables and constructs analytical performance/resource models with theoretical lower-bound guarantees, solved globally via NLP. Integrated with pragma semantic analysis and the Merlin compiler, and augmented by design-space pruning, the framework explores billion-scale configurations within seconds to minutes. Contribution/Results: Experimental evaluation shows kernel performance approaching hand-tuned implementations, resource estimation error <8%, and latency lower-bound error ≤12%.
Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.
This work proposes CuTe, a novel mathematical framework for representing and manipulating hierarchical tensor layouts through a layout algebra that supports operations such as concatenation, tiling, and inversion. Modern high-performance computing and deep learning rely heavily on specialized tensor instructions whose performance and correctness are critically dependent on intricate, hardware-specific data layouts. CuTe enables unified compile-time derivation, verification, and thread mapping of these layouts, significantly simplifying GPU kernel development while supporting expressive, general-purpose tensor transformations. The framework has been integrated into production systems including the NVIDIA CUTLASS library and the CuTe DSL, effectively bridging the gap between hardware constraints and software flexibility.
This work addresses the challenge of efficiently generating vector-length-agnostic (VLA) machine learning code for scalable vector instruction sets such as Arm SVE, where unknown vector lengths at compile time hinder traditional compilers. The authors present the first end-to-end VLA support in MLIR/IREE, introducing a vector-length-aware compact data layout and unifying dynamic tiling, operator fusion, and scalable vectorization within a single compilation framework. Evaluated on Arm CPUs, the generated SVE code achieves up to 1.45× speedup over IREE’s NEON implementation, outperforms multiple frameworks in the PyTorch ecosystem, and demonstrates strong scalability with increasing vector lengths in simulation, effectively balancing performance and hardware portability.
This work addresses the performance limitations of traditional pointer analysis by proposing a decoupled acceleration paradigm. Instead of tightly coupling simplification rules with the analysis itself—a common drawback of existing offline approaches—the method applies general-purpose, semantics-preserving compiler optimizations to the intermediate representation (IR) prior to analysis. This modular, analysis-agnostic strategy enhances efficiency without requiring modifications to the pointer analysis algorithm, thereby supporting seamless integration with diverse analyses. Empirical evaluation across multiple benchmark programs and three mainstream pointer analyses demonstrates speedups of up to 3.14× and memory reductions of up to 1.94×, all while largely preserving precision.
This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.