🤖 AI Summary
This work addresses the tight coupling between OpenMP semantics and fixed lowering strategies in existing compilers, which limits cross-hardware performance portability. Building upon the MLIR framework, we propose a modular parallel code generation approach that leverages domain-specific languages to declaratively specify lowering logic. By refactoring the lowering process into programmable components, our method decouples frontend semantics from backend targets, enabling explicit control over code outlining, data sharing, and runtime interfaces. Evaluations on the PolyBench/C-OMP benchmark suite demonstrate that this approach matches the performance of state-of-the-art compilers while introducing less than 0.7% code overhead. Furthermore, it reduces lowering code volume by 32% and 76% compared to Clang and GCC, respectively, and facilitates rapid adaptation to new runtimes.
📝 Abstract
The increasing diversity of parallel hardware challenges existing compilation flows. While OpenMP provides a portable abstraction for shared-memory parallelism, existing compilers tightly couple the frontend semantics with fixed lowering strategies. This design limits performance portability across different runtimes and architectures. In this paper, we present a modular approach to parallel code generation based on the Multi-Level Intermediate Representation (MLIR) framework. Our approach combines an OpenMP frontend with a domain-specific language (DSL) that specifies how OpenMP constructs are lowered for a target. We call this approach OpenMP meta-lowering, treating lowering as a programmable component rather than compiler-specific logic. This design provides explicit control over code outlining, data sharing, and runtime interfacing across diverse targets. We evaluate our approach on PolyBench/C-OMP across general-purpose and embedded multicore targets. Our results show that the proposed method matches the performance of state-of-the-art compiler toolchains with negligible code-size overhead (< 0.7%). Preserving OpenMP constructs until late lowering stages enables both general-purpose and OpenMP-specific MLIR optimizations. By decoupling lowering from compiler internals, our approach expresses the same OpenMP subset in 5,182 lines of code, ~32% fewer than the corresponding lowering code in Clang and ~76% fewer than in GCC, while supporting the pmsis runtime costs only a few specification lines, enabling rapid support for new runtimes without modifying the compiler.