Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism

📅 2025-05-16
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing speculative decoding (SD) suffers from substantial latency due to sequential execution of the draft and target models. This paper proposes SpecBranch, the first framework to achieve parallelization in speculative decoding—inspired by processor branch prediction. It introduces a rollback-aware branching mechanism and jointly optimizes adaptive draft length with hybrid confidence modeling, integrating implicit draft-model confidence and explicit target-model feature reuse. The method comprises four core components: branch-parallel decoding, hybrid draft generation, rollback-aware scheduling, and dynamic length control. Experiments demonstrate that SpecBranch achieves 1.8×–4.5× speedup over autoregressive decoding. Moreover, for low-alignment models, it reduces rollback tokens by 50%, significantly improving inference efficiency and deployment practicality.

Technology Category

Machine Learning: Mixture of Experts (MoE)Natural Language Processing: SummarizationCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Recently, speculative decoding (SD) has emerged as a promising technique to accelerate LLM inference by employing a small draft model to propose draft tokens in advance, and validating them in parallel with the large target model. However, the existing SD methods still remain fundamentally constrained by their serialized execution, which causes the mutual waiting bubbles between the draft and target models. To address this challenge, we draw inspiration from branch prediction in modern processors and propose a novel framework extbf{SpecBranch} to unlock branch parallelism in SD. Specifically, we first take an in-depth analysis of the potential of branch parallelism in SD, and recognize that the key challenge lies in the trade-offs between parallelization and token rollback. Based on the analysis, we strategically introduce parallel speculative branches to preemptively hedge against likely rejections. Meanwhile, to enhance parallelism, we jointly orchestrate adaptive draft lengths with a hybrid combination of the implicit draft model confidence and explicit reusing of target model features. Extensive experiments across various models and benchmarks show that SpecBranch achieves over extbf{1.8}$ imes sim$ extbf{4.5}$ imes$ speedups against the auto-regressive decoding and reduces rollback tokens by $ extbf{50}$% for poorly aligned models, realizing its applicability for real-world deployments.
Problem

Research questions and friction points this paper is trying to address.

Serialized execution causes mutual waiting bubbles in speculative decoding
Trade-offs between parallelization and token rollback limit acceleration potential
Poorly aligned models suffer from high token rejection rates during decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unlocking branch parallelism in speculative decoding
Using hybrid drafting with implicit and explicit features
Employing parallel branches to preemptively hedge rejections
🔎 Similar Papers
2023-12-18Neural Information Processing SystemsCitations: 52
Y
Yuhao Shen
Zhejiang University
J
Junyi Shen
National University of Singapore
Q
Quan Kong
Zhejiang University
T
Tianyu Liu
University of Science and Technology of China
Y
Yao Lu
National University of Singapore
C
Cong Wang
Zhejiang University