🤖 AI Summary
This study addresses the inefficiency of token-by-token generation in diffusion language models for long sequences by proposing the BLD framework. This method introduces a novel block-level denoising mechanism that compresses over a thousand consecutive tokens into a small number of block latent variables, combined with branched decoding to enable parallel local autoregressive generation. By integrating continuous diffusion, latent compression, and grouped conditional generation techniques, BLD effectively preserves both textual fluency and diversity while substantially accelerating long-text inference. Specifically, the proposed framework reduces computational costs by 80-fold and increases throughput by more than six times, offering a highly efficient solution for scaling diffusion-based language models to longer contexts.
📝 Abstract
Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than $80\times$ and increases throughput by more than $6\times$ relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than $6\times$ higher throughput and more than $4\times$ lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.