🤖 AI Summary
This study addresses the hardware adaptation challenges in distributed inference caused by the tight coupling of collective communication semantics, orchestration, and data paths. To this end, this work proposes a hierarchical decoupling framework comprising three layers: a top layer defining layouts and operations, a middle layer unifying orchestration logic via the SNAC protocol, and a bottom layer executing hardware-specific data movement using Atom primitives. By separating coordination logic from data paths, this architecture enables flexible customization and rapid integration of novel hardware mechanisms. The implementation optimizes GPU collective communication and integrates with the SGLang system. Experimental results demonstrate that the proposed approach reduces communication latency by 5.14× and improves bandwidth by 4.5×. Furthermore, LLM serving throughput increases by 1.37×, with online interactive performance improving by up to 2.85×.
📝 Abstract
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.