PipeDRAM: A Data-Transposition-Free In-DRAM Architecture with Hardware/Software Pipelining

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the transposition overhead caused by mismatches between horizontal and vertical data layouts in Processing-Unit-in-DRAM (PUD) architectures. To overcome this limitation, this work proposes PipeDRAM, a processing-in-memory architecture that eliminates runtime data layout conversion through deterministic bit-level reorganization. Furthermore, it introduces a hardware-software co-pipelining mechanism that exploits bit-level parallelism to overlap DRAM operations, thereby enabling horizontal in-memory computing without data transposition. Experimental results demonstrate that PipeDRAM achieves up to an 80.4× performance improvement and a 38× energy reduction with minimal hardware area overhead, providing an efficient architectural solution for processing-in-memory systems.
📝 Abstract
Processing-using-DRAM (PUD) architectures exploit the analog operational properties of DRAM to perform bulk bitwise Boolean and arithmetic operations inside memory arrays by organizing data in a vertical layout, where operand bits are stacked along DRAM columns. However, modern computing systems natively employ a horizontal data layout that preserves the cache line abstraction, leverages spatial locality in row buffers, and enables high memory throughput. This fundamental mismatch forces existing PUD architectures to frequently perform data layout transformations between horizontal and vertical formats, incurring significant performance, energy, and system integration overheads. Our goal is to eliminate data transposition overheads in PUD systems at low cost. To this end, we propose PipeDRAM, a PUD architecture that eliminates the need for runtime data layout transformation, enabling PUD operations directly over horizontally laid-out data. PipeDRAM's key ideas are to (i) deterministically reorganize bits inside each memory request to enable a PUD-friendly data placement within a DRAM array in a horizontal data layout, and (ii) employ a pipeline-based execution model that overlaps bit-dependent and bit-independent in-DRAM operations to exploit bit-level parallelism across the memory array. We compare PipeDRAM to different computing platforms. PipeDRAM provides (i) 11.8x, 11.8x, and 80.4x higher performance and (ii) 25.4x, 3.0x, and 38.0x lower energy consumption than three state-of-the-art PUD systems. PipeDRAM incurs low area cost on top of a DRAM chip (1.86%) and CPU die (0.05%). To enable further research on PUD systems, we open-source PipeDRAM at https://github.com/CMU-SAFARI/PipeDRAM.
Problem

Research questions and friction points this paper is trying to address.

Processing-using-DRAM
data transposition
data layout mismatch
in-memory computing
DRAM architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Processing-using-DRAM
Data-Transposition-Free
Pipelining
Bit-level Parallelism
Horizontal Data Layout
🔎 Similar Papers
2024-08-28IEEE Transactions on Computer-Aided Design of Integrated Circuits and SystemsCitations: 0