Fast and Memory Efficient Offload Training Framework with Hybrid XPU Computation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the GPU memory bottlenecks in large model training, as well as the low resource utilization and poor communication-computation overlap inherent in ZeRO-Offload. To tackle these challenges, this work proposes MemFerry, a framework leveraging Direct Host Access (DHA) technology. By employing a shadow model for unified memory abstraction alongside dynamic programming algorithms, MemFerry enables hybrid computation and dynamic offloading. Furthermore, it extends to multi-path transmission to optimize communication-computation overlap. Experimental results demonstrate that MemFerry achieves a 1.68× training speedup on a single GPU while supporting a 1.52× larger model scale. In multi-GPU settings, it yields over 28% acceleration and reduces iteration time by 20.7% on Huawei Cloud nodes.
📝 Abstract
With the ever-growing size of deep learning models, GPU memory is prone to being insufficient during training. A prominent approach is ZeRO-Offload, which moves the optimizer states to CPU memory and performs parameter update using CPU. However, the deficiencies of ZeRO-Offload include low GPU utilization, imperfect overlapping of communication and computation, and inflexible offloading. In this paper, we leverage Direct Host Access (DHA) on the GPU that can compute data in CPU memory, forming a novel hybrid on-GPU and DHA. We design and implement MemFerry consisting of an execution scheduler and a shadow model. The scheduler strategically chooses layers of parameters for DHA computation and transmits the remaining parameters to GPU memory simultaneously to shorten forward propagation time, and further loads DHA parameters to GPU memory to reduce backward propagation time. The shadow model presents a unified memory abstraction for the parameter partitions stored separately in GPU and CPU memories. To further reduce GPU memory usage, we present MemFerry along with its dynamic programming algorithm that offloads gradients to CPU memory via DHA. We further extend MemFerry to emerging scale-up domains with ScaleUp-MemFerry, which exploits otherwise underutilized accelerator interconnect bandwidth to assist host-to-accelerator data movement through adaptive multi-path transfer. Our experiments show that \system trains up to $1.68\times$ faster and MemFerry can train $1.52\times$ larger model compared to ZeRO-Offload on a single GPU, and increase training speed by at least $28.1\%$ when scaling to data parallelism on 8 GPUs. We further extend the design to a Huawei CloudMatrix384 scale-Up node with up to 8 NPUs, and our ScaleUp-MemFerry reduces the end-to-end iteration time by up to $20.7\%$ over DeepSpeed.
Problem

Research questions and friction points this paper is trying to address.

Offload Training
GPU Memory
ZeRO-Offload
Deep Learning
Hybrid XPU Computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offload Training
Direct Host Access
Hybrid XPU Computation
Memory Optimization
Scale-Up Interconnect
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyi Yao
College of Future Information Technology, Fudan University
Z
Zuning Liang
College of Future Information Technology, Fudan University
Yuedong Xu
Yuedong Xu
Professor, Fudan University
Machine Learning SystemsNetwork EconomicsNetwork ModelingMultimedia Networking
J
Jin Zhao
College of Computer Science and Artificial Intelligence, Fudan University
J
Jessie Hui Wang
Institute for Network Sciences and Cyberspace, Tsinghua University
Tong Li
Tong Li
Associate Professor, Renmin University of China. Chief Engineer, Huawei. PhD, Tsinghua University.
Computer NetworkingNetwork SecurityDistributed SystemsData Space