SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed attention parallelism layouts in adapting to dynamic workloads during large language model (LLM) serving. To overcome this, we propose a transparent switching mechanism that decouples KV cache from weight sharding. By introducing a Decoupled Ownership Parallelism (DOP) strategy to optimize GPU memory management, and integrating background state migration, batch-boundary handoffs, and load-aware scheduling algorithms, our approach eliminates the restart requirement of conventional inference engines, enabling seamless runtime transitions between parallelism layouts. Experimental evaluations demonstrate that the proposed system achieves a 1.3× to 1.73× throughput improvement on GPUs such as the B200, with switching overhead remaining below 0.51%.
📝 Abstract
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.
Problem

Research questions and friction points this paper is trying to address.

LLM serving
attention parallelism
dynamic workloads
layout switching
KV cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Attention Parallelism
Seamless Handoff
Decoupled Ownership Parallelism
KV Cache Decoupling
LLM Serving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Chuan Liu
Chuan Liu
University of Rochester
S
Shuoming Zhang
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Z
Zhicheng Li
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Q
Qianqi Sun
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
R
Ruiyuan Xu
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Q
Qiuchu Yu
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Xiyu Shi
Xiyu Shi
Institute for Digital Technologies, Loughborough University London
Speech signal processmobile and wireless communicationnetwork securityInternet of things
H
Huimin Cui
State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Jiacheng Zhao
Jiacheng Zhao
Institute of Computing Technology, Chinese Academy of Scienses
Parallel ComputingParallel CompilingComputer ArchitectureProgramming ModelDatacenter