DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical challenges in highly concurrent block-sparse speculative decoding, including verification padding, candidate rejection, and the incompatibility between variable-length prefixes and fixed computation graphs. To overcome these issues, this work proposes a systematic optimization framework. Methodologically, it introduces path-aware tiling, dynamic verification length allocation, and fixed-address workspaces to facilitate CUDA Graph reuse. Furthermore, a lightweight predictor is incorporated to enable adaptive verification without requiring confidence calibration or hardware profiling, thereby preserving draft model integrity. Experimental evaluations on the Qwen3 model series demonstrate that the proposed approach achieves throughput improvements of 43.9%–48.8% and reduces per-step decoding latency by 30.8%–52.5%.
📝 Abstract
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
block-diffusion
high concurrency
inference efficiency
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Block-Diffusion
Adaptive Verification
Dynamic Verify-Length
Graph Reuse
R
Rongjian Chen
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518052, China; and University of Chinese Academy of Sciences, Beijing 100049, China
Minxian Xu
Minxian Xu
Associate Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingMicroservicesLLM Inference
Z
Zhengxin Fang
Victoria University of Wellington, Wellington, New Zealand
Kejiang Ye
Kejiang Ye
Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingAI SystemsIndustrial Internet
C
Chengzhong Xu
Institute of AI and Brain Sciences, Department of Computer Science, University of Macau, Macau 999078, China