From Position Risks to Block Survival: Faster Generation for Diffusion Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prefix-condition mismatch and error asymmetry between parallel prediction and verification decoding in diffusion language models by proposing the BRISK-DLM framework. Methodologically, it introduces risk-reward weighted training to optimize proposal learning and designs a lightweight prefix-condition corrector that avoids additional backbone evaluation overhead. Furthermore, the framework integrates self-generated sequence training, dynamic positional priority ranking, and fused-execution inference to enhance verification efficiency. Experimental results demonstrate that BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task performance quality, thereby establishing a new performance boundary for accelerated decoding in diffusion language models.
📝 Abstract
Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Language Models
Parallel Decoding
Proposal-Verification
Token Generation
Throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Parallel Decoding
Risk-Reward Weighting
Prefix-Conditioned Corrector
Throughput Optimization
🔎 Similar Papers
No similar papers found.