🤖 AI Summary
This study addresses critical challenges in highly concurrent block-sparse speculative decoding, including verification padding, candidate rejection, and the incompatibility between variable-length prefixes and fixed computation graphs. To overcome these issues, this work proposes a systematic optimization framework. Methodologically, it introduces path-aware tiling, dynamic verification length allocation, and fixed-address workspaces to facilitate CUDA Graph reuse. Furthermore, a lightweight predictor is incorporated to enable adaptive verification without requiring confidence calibration or hardware profiling, thereby preserving draft model integrity. Experimental evaluations on the Qwen3 model series demonstrate that the proposed approach achieves throughput improvements of 43.9%–48.8% and reduces per-step decoding latency by 30.8%–52.5%.
📝 Abstract
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash