OpenMP Meta-Lowering: A Declarative Approach to Performance Portable Parallel Code Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tight coupling between OpenMP semantics and fixed lowering strategies in existing compilers, which limits cross-hardware performance portability. Building upon the MLIR framework, we propose a modular parallel code generation approach that leverages domain-specific languages to declaratively specify lowering logic. By refactoring the lowering process into programmable components, our method decouples frontend semantics from backend targets, enabling explicit control over code outlining, data sharing, and runtime interfaces. Evaluations on the PolyBench/C-OMP benchmark suite demonstrate that this approach matches the performance of state-of-the-art compilers while introducing less than 0.7% code overhead. Furthermore, it reduces lowering code volume by 32% and 76% compared to Clang and GCC, respectively, and facilitates rapid adaptation to new runtimes.
📝 Abstract
The increasing diversity of parallel hardware challenges existing compilation flows. While OpenMP provides a portable abstraction for shared-memory parallelism, existing compilers tightly couple the frontend semantics with fixed lowering strategies. This design limits performance portability across different runtimes and architectures. In this paper, we present a modular approach to parallel code generation based on the Multi-Level Intermediate Representation (MLIR) framework. Our approach combines an OpenMP frontend with a domain-specific language (DSL) that specifies how OpenMP constructs are lowered for a target. We call this approach OpenMP meta-lowering, treating lowering as a programmable component rather than compiler-specific logic. This design provides explicit control over code outlining, data sharing, and runtime interfacing across diverse targets. We evaluate our approach on PolyBench/C-OMP across general-purpose and embedded multicore targets. Our results show that the proposed method matches the performance of state-of-the-art compiler toolchains with negligible code-size overhead (< 0.7%). Preserving OpenMP constructs until late lowering stages enables both general-purpose and OpenMP-specific MLIR optimizations. By decoupling lowering from compiler internals, our approach expresses the same OpenMP subset in 5,182 lines of code, ~32% fewer than the corresponding lowering code in Clang and ~76% fewer than in GCC, while supporting the pmsis runtime costs only a few specification lines, enabling rapid support for new runtimes without modifying the compiler.
Problem

Research questions and friction points this paper is trying to address.

OpenMP
performance portability
parallel code generation
compiler lowering
hardware diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

OpenMP meta-lowering
MLIR
declarative DSL
performance portability
compiler decoupling
🔎 Similar Papers
No similar papers found.