PIM or CXL-PIM? Understanding Architectural Trade-offs Through Large-Scale Benchmarking

📅 2025-11-18
📈 Citations: 0
Influential: 0
📄 PDF

career value

253K/year
🤖 AI Summary
This paper systematically investigates the performance trade-offs between Processing-in-Memory (PIM) and CXL-based PIM (CXL-PIM) architectures. It addresses the fundamental tension: conventional PIM incurs high explicit data movement overhead, whereas CXL-PIM offers a unified address space but suffers from elevated memory access latency. To resolve this, the authors propose an end-to-end evaluation methodology that integrates empirical measurements from real PIM hardware with trace-driven CXL channel modeling, enabling large-scale benchmarking across mainstream workloads. Their analysis reveals, for the first time, that the amortization effect of unified addressing on interconnect latency is dynamic—varying with dataset size and execution phase—and can cause performance inversion between the two architectures. Building on this, they quantitatively characterize the boundary conditions defining the performance crossover points. The findings provide principled, quantifiable guidance for near-memory system design and uncover a novel architectural design space.

Technology Category

Application Category

📝 Abstract
Processing-in-memory (PIM) reduces data movement by executing near memory, but our large-scale characterization on real PIM hardware shows that end-to-end performance is often limited by disjoint host and device address spaces that force explicit staging transfers. In contrast, CXL-PIM provides a unified address space and cache-coherent access at the cost of higher access latency. These opposing interface models create workload-dependent tradeoffs that are not captured by small-scale studies. This work presents a side-by-side, large-scale comparison of PIM and CXL-PIM using measurements from real PIM hardware and trace-driven CXL modeling. We identify when unified-address access amortizes link latency enough to overcome transfer bottlenecks, and when tightly coupled PIM remains preferable. Our results reveal phase- and dataset-size regimes in which the relative ranking between the two architectures reverses, offering practical guidance for future near-memory system design.
Problem

Research questions and friction points this paper is trying to address.

Comparing architectural trade-offs between PIM and CXL-PIM memory systems
Identifying workload conditions favoring unified versus disjoint address spaces
Determining when each architecture overcomes data transfer bottlenecks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Processing-in-memory reduces data movement
CXL-PIM provides unified address space access
Large-scale benchmarking reveals workload-dependent tradeoffs
🔎 Similar Papers
No similar papers found.
I
I-Ting Lee
Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan
B
Bao-Kai Wang
Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan
Liang-Chi Chen
Liang-Chi Chen
Department of Computer Science and Information Engineering, National Taiwan University, Taipei, Taiwan
W
Wen Sheng Lim
Department of Computer Science and Information Engineering, National Taiwan University, Taipei, Taiwan
D
Da-Wei Chang
Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan
Yu-Ming Chang
Yu-Ming Chang
Wolley, Taipei, Taiwan
C
Chieng-Chung Ho
Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan