Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism

📅 2026-01-09
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional parallel speculative decoding, which is constrained by a theoretical speedup ceiling and prone to high computational waste and pipeline stalls due to early prediction errors. The authors propose Double, a novel framework that introduces, for the first time, a synchronized dual-retrieval mechanism combining iterative retrieval-based speculation with authoritative multi-token guidance from the target model. This approach breaks through the acceleration bottleneck without requiring any additional training, effectively mitigating the trade-off between accuracy and efficiency while achieving lossless speedup. Experimental results demonstrate that Double achieves 5.3× and 2.8× acceleration on LLaMA3-70B and Qwen3-32B, respectively, significantly outperforming existing methods such as EAGLE-3 that rely heavily on extensive training.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to SearchNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Parallel Speculative Decoding (PSD) accelerates traditional Speculative Decoding (SD) by overlapping draft generation with verification. However, it remains hampered by two fundamental challenges: (1) a theoretical speedup ceiling dictated by the speed ratio between the draft and target models, and (2) high computational waste and pipeline stall due to mid-sequence token rejections of early errors. To address these limitations, we introduce \textsc{Double} (Double Retrieval Speculative Parallelism). By bridging the gap between SD and PSD, our framework resolves the Retrieval \emph{Precision-Efficiency Dilemma} through a novel synchronous mechanism. Specifically, we enable the draft model to execute iterative retrieval speculations to break the theoretical speedup limits; to alleviate rejections without rollback, the target model performs authoritative retrieval to generate multi-token guidance. \textsc{Double} is entirely training-free and lossless. Extensive experiments demonstrate state-of-the-art speedup of $\textbf{5.3}\times$ on LLaMA3.3-70B and $\textbf{2.8}\times$ on Qwen3-32B, significantly outperforming the advanced method EAGLE-3 that requires extensive model training.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Acceleration Limit
Computational Waste
Pipeline Stall
Speedup Ceiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Parallel Speculative Decoding
Double Retrieval
Training-Free Acceleration
Token Rejection Mitigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuhao Shen
Zhejiang University
Tianyu Liu
Tianyu Liu
Hongkong University of Science and Technology
J
Junyi Shen
National University of Singapore
J
Jinyang Wu
Tsinghua University
Q
Quan Kong
Zhejiang University
L
Li Huan
Zhejiang University
Cong Wang
Cong Wang
Zhejiang University
LLM Safety/Efficiency