The Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that mandatory sequential constraints in imperative code hinder parallelizing compilers from proving and eliminating redundant dependencies. To overcome this, it proposes the Canonical Parallel Form (CPF), which constructs deterministic, device-neutral program states through tensor semantics lifting. This approach integrates a three-tiered hierarchical dependency analysis—encompassing syntactic, affine, and SMT-based techniques—to achieve effective redundancy elimination. The resulting framework enables device-agnostic, high-performance parallel compilation while significantly reducing the token overhead associated with agent-driven optimization. Experimental evaluations on AMD MI300A demonstrate up to 25.4× speedup over Numba, outperforming established tools such as DaCe and Pluto, alongside a 2.72× reduction in token consumption.
📝 Abstract
Imperative code fixes an execution order the computation does not require, and a parallelizing compiler must prove which parts of that order it can remove. We introduce the Canonical Parallel Form (CPF), a device-neutral program state from which every ordering constraint our analyses prove unnecessary has been removed. CPF is reached by output-preserving normalization, by lifting semantic operations such as tensor contractions, and by deriving parallelism in three levels ordered by decidability: syntactic subscript tests, exact affine dependence tests over integer sets, and SMT queries over non-linear integer arithmetic, which also admit parallelism guarded behind a runtime check. Heuristics can then specialize the canonical form for each architecture. Across 248 loop-level reasoning kernels on an AMD MI300A, CPF reaches 4.4x over Numba on its 24 Zen 4 cores and 25.4x on its CDNA 3 GPU. Against the other auto-parallelizing optimizers, CPF is 2.9x faster on the CPU and 8.7x on the GPU than DaCe's own auto-parallelizer, 1.9x faster than Pluto, and 1.5x faster than PPCG on the affine subset of the kernels. Because the pipeline is deterministic, CPF also serves as an agent's starting source, cutting the token cost per kernel by up to a factor of 2.72x while leaving the achieved speed-up unchanged, since the coding agents reason less about parallelism.
Problem

Research questions and friction points this paper is trying to address.

parallelizing compilers
agentic optimizers
imperative code
execution order
canonical parallel form
Innovation

Methods, ideas, or system contributions that make the work stand out.

Canonical Parallel Form
Parallelizing Compilers
Agentic Optimizers
Affine Dependence Analysis
SMT Solving
🔎 Similar Papers
No similar papers found.