Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff

📅 2025-02-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing speculative decoding (SD) mandates strict distributional equivalence between the draft and target models, limiting acceleration potential and depriving users of flexibility in trading off quality against latency. This paper proposes Fuzzy Speculative Decoding (FSD), the first SD framework to incorporate a distributional tolerance mechanism: it quantifies and dynamically controls the divergence—measured via KL divergence—between draft and target model distributions, enabling user-controllable precision–latency trade-offs without modifying model weights or altering existing SD architectures. We theoretically prove that strict distributional equivalence is not necessary for maintaining performance. Experiments across multiple benchmarks show that FSD achieves speedups exceeding 5 tokens/s over standard SD with only ~2% accuracy degradation; in several scenarios, it maintains accuracy while accelerating inference by over 2 tokens/s.

Technology Category

Search and Optimization: Distributed SearchMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Speculative Decoding (SD) enforces strict distributional equivalence to the target model, limiting potential speed ups as distributions of near-equivalence achieve comparable outcomes in many cases. Furthermore, enforcing distributional equivalence means that users are unable to trade deviations from the target model distribution for further inference speed gains. To address these limitations, we introduce Fuzzy Speculative Decoding (FSD) - a decoding algorithm that generalizes SD by accepting candidate tokens purely based on the divergences between the target and draft model distributions. By allowing for controlled divergence from the target model, FSD enables users to flexibly trade generation quality for inference speed. Across several benchmarks, our method is able to achieve significant runtime improvements of over 5 tokens per second faster than SD at only an approximate 2% absolute reduction in benchmark accuracy. In many cases, FSD is even able to match SD benchmark accuracy at over 2 tokens per second faster, demonstrating that distributional equivalence is not necessary to maintain target model performance.
Problem

Research questions and friction points this paper is trying to address.

Enables flexible tradeoff between generation quality and inference speed.
Introduces Fuzzy Speculative Decoding to allow controlled divergence from target model.
Achieves significant runtime improvements with minimal accuracy reduction.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces Fuzzy Speculative Decoding (FSD)
Allows controlled divergence for speed gains
Achieves significant runtime improvements
🔎 Similar Papers
No similar papers found.