ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing distributed LLM serving simulators lack closed-loop execution and hardware-free profiling for timing prediction. This work proposes a closed-loop simulation framework that integrates specification-driven analytical timing modeling with a stateful serving loop. It leverages iSTAGE to generate profile-free traces and decouples component ownership to support cross-platform portability. Compatible with the vLLM interface, the framework enables unmodified benchmarks to run directly while accurately capturing scheduling, queuing, and KV cache feedback. Experiments demonstrate steady-state throughput and multi-turn performance errors of only 3.6% and 9.9%, respectively. Furthermore, the study reveals novel mechanisms, including an inversion in the HBM bandwidth-capacity trade-off and a concurrency-induced bottleneck shift from memory constraints to scheduling overhead.
📝 Abstract
Evaluating distributed LLM serving designs on physical clusters is costly. Yet existing simulators provide only subsets of the capabilities needed for realistic design exploration: stateful closed-loop execution, timing prediction without profiling target hardware, and direct execution of unmodified serving benchmarks. We present ServeTwin, a closed-loop simulator that couples specification-driven analytical timing with a stateful serving loop. This coupling captures feedback among request completion, scheduling, queue state, and KV-cache evolution, including prefill-decode disaggregation and multi-turn agent workloads. ServeTwin avoids target-hardware operator profiling through iSTAGE, an analytical trace generator that derives per-iteration traces from model and engine specifications while representing ragged batches and discrete engine mechanisms such as CUDA-graph batch padding. It further decomposes execution time into component-owned throughput, scheduler, and runtime costs. This ownership lets unaffected parameters transfer across platforms and confines recalibration to changed hardware or software components. ServeTwin implements a vLLM-compatible interface and runs unmodified serving benchmarks. Against real deployments, it reproduces InferenceX's steady-state throughput-interactivity frontier with a 3.6% mean error and predicts LMBenchmark's multi-turn performance with a 9.9% error while tracking KV-cache evolution. Pre-silicon sweeps reveal that the preferred HBM bandwidth-capacity tradeoff reverses across workload states and that increasing concurrency shifts the bottleneck from memory bandwidth to scheduler and runtime overheads. Together, these capabilities enable practical exploration of distributed LLM serving systems before target hardware is available.
Problem

Research questions and friction points this paper is trying to address.

Distributed LLM serving
Simulation
Architecture exploration
Benchmark validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributed LLM Serving
Closed-loop Simulation
Analytical Timing Model
Hardware-agnostic Profiling
KV-cache Evolution
🔎 Similar Papers
No similar papers found.
S
Sungjoon Park
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
C
Changue Jung
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
K
Kyungno Joo
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
M
Mincheol Kang
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
J
Jaehyung Ahn
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
S
Sehwan Lee
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea
S
Sangjoon Kim
Samsung Advanced Institute of Technology, Suwon-si, Republic of Korea