NebulaSD: Many-for-Many Speculative Decoding

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency in speculative decoding for large language models caused by conflicting batch configurations between drafting and verification phases, as well as rigidly coupled resource allocation. To overcome these limitations, this work proposes a many-to-many speculative decoding system that independently schedules resource pools to enable global capacity sharing. Furthermore, it introduces a dynamic reallocation mechanism without fixed locality, which integrates worker-triggered batch reconstruction with asynchronous KV cache management to eliminate migration stalls. Experimental evaluations on a four-GPU deployment demonstrate that the proposed approach improves request throughput by 50.4% compared to the baseline and by 72.6% relative to co-located execution, thereby significantly enhancing GPU utilization.
📝 Abstract
Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Drafting and verification, however, exhibit different service characteristics and favor different batch configurations, making fixed draft-target coupling inefficient under concurrent workloads. Existing distributed designs can physically separate the two stages, but often retain request or batch affinities that prevent their capacities from being shared globally. We present NebulaSD, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools. Such dynamic reassignment removes fixed worker locality, requiring request states to be made available at newly selected workers without introducing migration stalls. NebulaSD addresses this challenge through worker-triggered batch reconstruction and asynchronous KV-state preparation overlapped with model execution. We evaluate NebulaSD from both system and scaling perspectives, showing that dynamic pooling improves request-round processing rate by 50.4% over a physically disaggregated baseline and 72.6% over co-located execution on a four-GPU deployment while substantially increasing effective GPU utilization. Profile-driven simulations further show approximately proportional compute-side capacity scaling under idealized state movement.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
large language model inference
resource pooling
batch scheduling
state migration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Many-for-Many Scheduling
Dynamic Batch Reconstruction
Asynchronous KV-state Preparation
Resource Pooling
🔎 Similar Papers
No similar papers found.