D-Loop: Looped Diffusion Drafting for Speculative Decoding

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the repetition traps and degraded draft quality in block diffusion speculative decoding caused by neglecting preceding tokens. To overcome these limitations, we propose D-Loop, which introduces the first intra-block causal conditioning mechanism. By employing recurrent diffusion drafting that reuses a single backbone network for iterative prediction, D-Loop models causal dependencies during multi-token parallel generation without requiring additional components. Furthermore, it incorporates complementary prefix-suffix training objectives to optimize generation quality. Experimental results demonstrate that D-Loop significantly outperforms existing methods such as DFlash and DSpark across mathematical, coding, and dialogue benchmarks on the Qwen3 model series.
📝 Abstract
Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
block diffusion
repetition trap
draft quality
marginal distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Block Diffusion
Intra-block Causal Conditioning
Parameter Sharing
Prefix-Suffix Objective