SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of resource contention, transmission redundancy, and limited strong scaling in multi-GPU host-resident fine-tuning by proposing a synchronous data-parallel runtime system. The system maintains a single authoritative host state to decouple communication from memory layout, enabling pipelined execution of parameter delivery, gradient aggregation, and CPU-based updates. Furthermore, an analytical step-time model is introduced to guide scheduling optimization, facilitating efficient pipelining across ranks and chunks. Experimental results demonstrate that the proposed approach achieves a 1.46× to 2.64× throughput improvement, supports million-token batch sizes and ultra-long sequence training, and surpasses the peak performance of FSDP2.
📝 Abstract
Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.
Problem

Research questions and friction points this paper is trying to address.

LLM fine-tuning
host-resident
data parallelism
multi-GPU
memory bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Host-resident LLM fine-tuning
Synchronous data parallelism
Pipeline scheduling
Analytical step-time model
Memory-efficient training
R
Ruijia Yang
The Hong Kong University of Science and Technology (Guangzhou)
S
Shiyuan Lin
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yulong Ao
Beijing Academy of Artificial Intelligence
Zhiyu Li
Zhiyu Li
Tianjin University
Robust controlattitude control
Y
Yingli Zhao
Beijing Academy of Artificial Intelligence
X
Xianduo Li
Beijing Academy of Artificial Intelligence
Y
Yonghua Lin
Beijing Academy of Artificial Intelligence
Zeyi Wen
Zeyi Wen
Assistant Professor at HKUST(Guangzhou)
Efficient LLMsMLSysHPOHPC