Fast Recovery for LLM Serving via Decoupled Device Memory Lifetime in Dynamo

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM服务中因硬件或软件故障导致的长时间恢复问题,本文提出通过解耦设备内存生命周期来实现快速恢复的方法。
📝 Abstract
Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure can disrupt an entire replica. Recovery requires reinitializing the engine, taking minutes even when weights and compilation artifacts are cached. Production deployments overprovision serving capacity to mask this window. We argue that the dominant cost is loss of ready serving capacity, not request progress, so recovery should preserve initialized engine state rather than reconstruct it. We present fast recovery for Dynamo based on this principle. Snapshots capture an initialized engine once and restore it instead of reinitializing it. Analysis of 18 weeks of failures from the Dynamo cluster shows that most failures are device-preserving: the engine process fails while the GPU and its resident allocations remain intact. Our key insight is that independent engine processes can reuse the same GPU-resident state while keeping mutable execution state private. The GPU Memory Service (GMS) decouples device-memory ownership from engine processes, enabling engines to share and reattach surviving allocations without copying them. GMS preserves model weights and shares them read-only between replacement and Shadow Engines, avoiding weight reloads. A second initialized runtime on the same GPUs reduces recovery to promotion. Across four models on vLLM and SGLang, these mechanisms recover a failed replica in under 7 seconds, 13-29 times faster than a warm restart, using a fixed 4-8 GiB of device memory per GPU independent of model size. Replaying the production trace, we estimate they would reclaim 79% of GPU-hours lost to recovery.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Inference
Failure Recovery
Device Memory
Engine Initialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fast Recovery
GPU Memory Service (GMS)
Decoupled Device-Memory Ownership
Initialized Engine State Reuse
Model Weight Sharing
🔎 Similar Papers
S
Schwinn Saereesitthipitak
NVIDIA
M
Mohammed Abdulwahhab
NVIDIA
H
Hannah Zhang
NVIDIA
D
Dan Feigin
NVIDIA
N
Neelay Shah
NVIDIA
M
Maksim Khadkevich
NVIDIA
I
Itay Neeman
NVIDIA
V
Vikram Sharma Mailthody
NVIDIA
Wen-mei W. Hwu
Wen-mei W. Hwu
Senior Distinguished Research Scientist, NVIDIA; Professor and Sanders-AMD Chair of Electrical and
Computer ArchitectureCompilerParallel ComputingCognitive Computing Systems