🤖 AI Summary
Existing speculative decoding (SD) suffers from substantial latency due to sequential execution of the draft and target models. This paper proposes SpecBranch, the first framework to achieve parallelization in speculative decoding—inspired by processor branch prediction. It introduces a rollback-aware branching mechanism and jointly optimizes adaptive draft length with hybrid confidence modeling, integrating implicit draft-model confidence and explicit target-model feature reuse. The method comprises four core components: branch-parallel decoding, hybrid draft generation, rollback-aware scheduling, and dynamic length control. Experiments demonstrate that SpecBranch achieves 1.8×–4.5× speedup over autoregressive decoding. Moreover, for low-alignment models, it reduces rollback tokens by 50%, significantly improving inference efficiency and deployment practicality.
📝 Abstract
Recently, speculative decoding (SD) has emerged as a promising technique to accelerate LLM inference by employing a small draft model to propose draft tokens in advance, and validating them in parallel with the large target model. However, the existing SD methods still remain fundamentally constrained by their serialized execution, which causes the mutual waiting bubbles between the draft and target models. To address this challenge, we draw inspiration from branch prediction in modern processors and propose a novel framework extbf{SpecBranch} to unlock branch parallelism in SD. Specifically, we first take an in-depth analysis of the potential of branch parallelism in SD, and recognize that the key challenge lies in the trade-offs between parallelization and token rollback. Based on the analysis, we strategically introduce parallel speculative branches to preemptively hedge against likely rejections. Meanwhile, to enhance parallelism, we jointly orchestrate adaptive draft lengths with a hybrid combination of the implicit draft model confidence and explicit reusing of target model features. Extensive experiments across various models and benchmarks show that SpecBranch achieves over extbf{1.8}$ imes sim$ extbf{4.5}$ imes$ speedups against the auto-regressive decoding and reduces rollback tokens by $ extbf{50}$% for poorly aligned models, realizing its applicability for real-world deployments.