Score
Designs and implements code transformations that reorder nested loops by swapping loop nesting levels to change iteration order. Analyzes dependence constraints and expected performance effects such as memory locality, vectorization and parallelism to ensure the interchange is legal and beneficial, and integrates the transform into compiler or source-to-source optimization passes.
Poor generalizability of auto-schedulers across languages and projects arises from the structural diversity of nested loops, hindering consistent performance optimization. Method: This paper proposes *prior loop nest normalization*: before scheduling, it analyzes memory access patterns to determine loop equivalence and applies canonicalization rewrite rules to map heterogeneous nested loops onto a unified normalized form—departing from conventional post-hoc optimization paradigms. Contribution/Results: The lightweight frontend framework integrates seamlessly with schedulers such as Polly and Tiramisu. Evaluated on 15 multi-language benchmarks, it achieves 21.13× average speedup over C compilers—outperforming Polly and Tiramisu by 2.31× and 2.89×, respectively—and delivers 9.04×, 3.92×, and 1.47× improvements over NumPy, Numba, and DaCe. On the ECMWF CLOUDSC weather simulation code, it yields a 10% runtime reduction.
Existing polyhedral compilers suffer from limited support for deep affine transformations and rely on oversimplified program assumptions—such as rectangular iteration domains and single-level loop nests—hindering automatic selection of high-benefit schedules and impairing generality. This paper presents the first deep learning–driven clustering auto-scheduler designed for large-scale affine transformation spaces and complex program structures, including non-rectangular iteration domains and multi-level nested loops. Our approach integrates deep learning–based cost modeling, clustering-aware dependence analysis, multi-stage transformation sequence search, and iteration domain normalization with feature encoding. Evaluated on the PolyBench benchmark suite, our scheduler achieves geometric mean speedups of 1.84× over Tiramisu and 1.42× over Pluto. It significantly enhances optimization capability and practical applicability for complex, real-world programs.
This work addresses the challenge of efficiently verifying structural optimizations—such as loop unswitching and full loop unrolling—in verified compilers using small-step semantics, which struggles to precisely capture divergence and global control flow. The authors propose a novel hybrid approach that combines small-step and big-step semantics: small-step semantics is employed for local transformations, while coinductive big-step semantics accurately models divergent behaviors and handles structural transformations. An abstract behavioral semantics unifies the interfaces of both styles. This method enables, for the first time, the seamless integration of big-step semantics into CompCert’s predominantly small-step verification framework, achieving end-to-end formal verification of complex loop optimizations without altering the top-level semantic preservation theorem.
This work addresses the problem of dynamically maintaining loop nesting forests in reducible control flow graphs under edge insertions and deletions. It presents the first fully dynamic algorithm that locally updates the depth-first search (DFS) tree upon graph modifications, thereby avoiding costly global recomputation. The approach leverages dynamic DFS maintenance and incremental graph algorithms, supported by formal invariants and a rigorous correctness proof, to efficiently support real-time updates of loop structures and enable rapid derivation of dominance information. By providing a practical dynamic abstraction for compiler optimizations and program analysis, this study significantly enhances the responsiveness of modern compilation pipelines to changes in control flow.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
This work addresses the limitations of fixed branch ordering in function merging, which constrains optimization potential, and investigates the computational complexity when branch reordering is permitted—a setting that renders the problem NP-hard. For the first time, the problem is systematically studied through the lens of parameterized complexity, formulated as a parameterized sequence alignment task. The analysis reveals that the branching factor $b$ and nesting depth $d$ (or an alternative depth measure $d_2$) critically govern tractability. Leveraging this insight, we design fixed-parameter tractable (FPT) algorithms with running times $2^{O(bd)} n^2$ and $2^{O(bd_2)} n^7$, respectively. Moreover, we establish that the problem remains NP-hard even when certain parameters are fixed, thereby delineating a fine-grained boundary of its computational complexity.
This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.
Existing approaches struggle to accurately extract idempotent slices from general control flow graphs, limiting their applicability in code optimization. This work formally defines idempotent backward slicing for the first time and presents a sound and efficient slicing algorithm based on the Gated Static Single Assignment (GSA) representation, integrating precise control-flow and data-flow analyses. The proposed method enables merging of non-contiguous instructions across basic blocks and even across function boundaries, thereby overcoming the locality constraints inherent in traditional slicing techniques. Experimental evaluation demonstrates that, when applied to LLVM-optimized benchmark programs, the technique achieves up to a 7.24% reduction in code size.
This work proposes a novel approach to program transformation—such as compiler optimizations—by systematically integrating non-determinism and partially defined operations from functional logic programming. Leveraging the Curry language and its FlatCurry intermediate representation, the method expresses transformation rules in a concise, declarative style through pattern matching and non-deterministic computation. Traditional approaches often involve intricate manipulations of intermediate representations like abstract syntax trees, leading to implementations that are both cumbersome and error-prone. In contrast, the proposed technique enhances expressiveness and readability while maintaining practical feasibility. Experimental evaluation demonstrates that this approach not only preserves code clarity but also achieves competitive performance in real-world transformation tools, offering a compelling alternative for implementing reliable and maintainable program transformations.
本文解决了细粒度代码重构自动化问题,通过形式化五种Move Statement重构方法,并结合现有技术实现更细粒度的表达式移动。