Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of automated Processing-in-Memory (PIM) offloading support in deep learning frameworks and the associated data movement bottlenecks by proposing a compiler-based automated host-PIM placement strategy. Leveraging MLIR and PyTorch compiler technologies, this method overcomes the limitations of fixed operator lists by dynamically profiling the computational intensity and memory access characteristics of loop nests generated through progressive loop order lowering. This enables profile-guided optimization (PGO)-driven automatic offloading decisions. Evaluated across diverse PIM configurations, the proposed approach achieves speedups of 5.1ร— and 3.6ร— over CPU-only execution for GPT-J-6B and LLaMA-7B, respectively.
๐Ÿ“ Abstract
Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. However, PyTorch and other DL frameworks lack compiler support for making this decision on the code they lower: existing offloading frameworks target hand-written C/C++ programs, while those that address DL fix the candidate set to a list of operator types before lowering. We present Torch-PIM, a compiler framework that uses profile-guided optimization (PGO) to decide host-versus-PIM placement over the loop nests that progressive lowering materializes. Every parallel loop nest the pipeline emits enters the candidate space, and each is assessed in two stages: the amount of work it carries, and its memory boundedness. Every quantity the assessment consumes is profiled on the host or obtained from the multi-level intermediate representation (MLIR) of the code. Across PIM configurations of 32 to 128 cores, Torch-PIM's offloading decisions yield speedups of up to 8.6x on tensor operators, 2.9x on MLP, 4.4x on Attention, 5.1x on GPT-J-6B, and 3.6x on LLaMA-7B over CPU-only execution.
Problem

Research questions and friction points this paper is trying to address.

Processing-in-Memory
Deep Learning
Compiler Offloading
PyTorch
Data Movement Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Processing-in-Memory
Profile-Guided Optimization
Compiler Framework
MLIR
Deep Learning Offloading
H
Heeeon Lee
Yonsei University, Republic of Korea
Hyunwoo Nam
Hyunwoo Nam
Yonsei University, Republic of Korea
J
Junyong Heo
Yonsei University, Republic of Korea
H
Hyunmo Sung
Yonsei University, Republic of Korea
J
Jay Hwan Lee
Yonsei University, Republic of Korea
Y
Yeonsoo Kim
Yonsei University, Republic of Korea
S
Seongho Jeong
Yonsei University, Republic of Korea
S
Shinhyung Yang
Kiel University, Kiel, Germany
B
Bernd Burgstaller
Yonsei University, Republic of Korea