TRANSIT: Transparent Scale-in for Multi-Node LLM Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of GPU memory capacity and poor scalability in multi-node large model training by proposing Scale-in, an architecture based on user-space transparent interception. The proposed method leverages CPU DRAM to extend GPU memory, achieving efficient data transfer through a zero-copy data path and RoCE network optimization. Furthermore, it seamlessly integrates with distributed parallel strategies, enabling training without modifying the underlying system. Experimental results demonstrate that Scale-in improves throughput by 42%–68% over state-of-the-art methods. Notably, it maintains 90% training efficiency while reducing GPU usage by 50% and achieves a 35% speedup in communication-bound scenarios, significantly lowering hardware requirements and network overhead.
📝 Abstract
TRANSIT is a transparent scale-in framework to enable multi-node model training on fewer GPUs while maintaining training efficiency by transparently leveraging CPU DRAM as an extension of GPU memory during distributed training. It achieves this through a user-space interposition layer, requiring no modifications to the application, training framework, cluster scheduler, device driver, or operating system. Furthermore, TRANSIT achieves higher efficiency by leveraging a zero-copy data path for CPU-GPU transfers. We evaluate TRANSIT on dense and MoE models across scales up to 64 NVIDIA H100 GPUs and multiple parallelism configurations over a RoCE network. Our evaluation shows that TRANSIT can: (a) outperform state-of-the-art framework-managed offloading techniques, achieving up to 68%, 59%, and 42% higher per-GPU throughput than TorchTitan, ZeRO-Offload, and ZeRO-Infinity, respectively, (b) enables training with 50% fewer GPUs while maintaining over 90% of baseline per-GPU throughput, (c) lower per-node network traffic by up to 33%, and (d) improve per-GPU throughput by up to 35% in communication-bound settings.
Problem

Research questions and friction points this paper is trying to address.

multi-node LLM training
scale-in
memory offloading
distributed training
GPU efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transparent Scale-in
User-space Interposition
Zero-copy Data Path
Multi-node LLM Training
Memory Offloading
🔎 Similar Papers