DEdit: Iterative Draft Editing for Speculative Decoding

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in diffusion-based draft generation where independent predictions readily trigger error cascades, leading to the rejection of valid tokens. To overcome this limitation, we propose DEdit, a framework that introduces an iterative editing mechanism based on bidirectional attention to correct draft errors and thereby extend the accepted prefix. Furthermore, we design ProposalMix, a training strategy that blends model confidence with ground truth to rectify erroneous predictions while preserving correct ones, effectively transcending the constraints of conventional parallel demasking. Experiments conducted with Qwen3 demonstrate that DEdit achieves the highest macro-average token acceptance rate across multiple benchmarks, yielding a 5.72× to 5.97× speedup over autoregressive generation.
📝 Abstract
Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
diffusion-based drafter
independent token prediction
prefix verification
draft error propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Diffusion-based Drafter
Iterative Draft Editing
ProposalMix
Bidirectional Context
🔎 Similar Papers
No similar papers found.