π€ AI Summary
This study addresses the high computational cost and slow generation speed of diffusion large language models during reward-guided reasoning by proposing an adaptive hybrid decoding strategy. The method integrates KV caching with sparse attention recomputation to achieve parallel amortization, while introducing a confidence-based dynamic deferral mechanism to optimize autoregressive decoding efficiency. Experimental results demonstrate that this strategy achieves up to 4.4Γ inference acceleration without compromising generation quality, effectively resolving a critical bottleneck in the efficient generation of diffusion language models.
π Abstract
Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.