Compiling Semi-Ring Dynamic Programming to Tier-Aware 3D-DRAM Processing-in-Memory

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in 3D processing-in-memory (PIM) dynamic programming, which relies on hand-written kernels and manual optimization of data layout and communication. To overcome this, we propose GenMLIR, a compiler framework that pioneers semiring grid updates as a first-class intermediate representation (IR) abstraction. Built upon the MLIR infrastructure, it designs four GenDRAM-aware passes to enable fully automated tiling, placement, communication optimization, and code generation for semiring dynamic programming on hierarchy-aware 3D PIM architectures. Experimental results demonstrate that the proposed framework achieves up to a 16.7× performance improvement, approaching the theoretical upper bound, while reducing development effort from hundreds of lines of code to merely 4–9 lines.
📝 Abstract
Processing-in-Memory (PIM) on monolithic 3D (M3D) DRAM is a promising answer to the memory wall for data-intensive dynamic programming (DP), yet extracting its performance today demands hand-written kernels: the programmer must pick a tile size, place data across non-uniform-latency memory tiers, partition work between heterogeneous processing units, and insert the right broadcasts and queues by hand. We present GenMLIR, an MLIR compiler that automates this mapping for tier-aware 3D PIM. GenMLIR encodes the semi-ring generalized grid update, the algebraic form shared by all-pairs shortest path (APSP) and sequence alignment, as a first-class IR abstraction, and lowers it through four GenDRAM-aware pass groups that derive blocked tiling, tier-aware placement, tile-to-PU assignment, and explicit communication. We also characterize precisely which DP recurrences the abstraction admits and which it does not. On the GenDRAM architecture, GenMLIR-generated code runs up to 5.8 (APSP) and 16.7 (alignment) faster than a lowering without compiler support, and 1.2-1.4 faster than a standard affine-tiling PIM compiler we also implement, reaching 90-100% of an achievable-performance bound. It expresses these workloads in 4-9 lines instead of 300-500 and compiles in a negligible fraction of runtime.
Problem

Research questions and friction points this paper is trying to address.

Processing-in-Memory
Dynamic Programming
3D-DRAM
Memory Wall
Compiler Automation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Processing-in-Memory
MLIR Compiler
Semi-Ring Dynamic Programming
3D-DRAM
Tier-Aware Placement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mahbod Afarin
University of California San Diego
T
Tsung-Han Lu
University of California San Diego
Tajana Rosing
Tajana Rosing
Distinguished Professor, UCSD
computer architecturecyber-physical systemssystem energy efficiency