Score
Designs, implements, and evaluates probabilistic sequence models and decoding algorithms that generate or recover natural-language token sequences by applying iterative diffusion/denoising processes instead of autoregressive factorization. This includes training objectives, architectures and sampling schedules for masked-token denoising, bidirectional/infill-capable decoding, and diffusion-based text generation, together with analyses of convergence, sampling speed, and tradeoffs versus autoregressive approaches.
Diffusion language models (DLMs) hold promise for high-quality, parallelizable text generation but suffer from inefficiency, weak long-sequence modeling, and substantial infrastructure overhead. To address these challenges, this work conducts a systematic literature review and establishes the first unified technical taxonomy for DLMs, clarifying their theoretical relationships with autoregressive and masked language models. We comprehensively analyze key components—including pretraining and post-training paradigms, iterative denoising mechanisms, cache-augmented parallel decoding, generation quality optimization, and multimodal fusion architectures. Empirical evaluation demonstrates that our integrated methodology achieves multi-fold inference speedup while maintaining generation quality competitive with state-of-the-art autoregressive models. This work provides a foundational framework and practical guidelines for advancing DLM research, engineering optimization, and cross-modal extension.
To address the limitations of autoregressive language models—including poor parallelizability, weak fine-grained controllability, and inadequate response awareness—this survey systematically reviews recent advances in discrete diffusion language models (dLLMs) and multimodal diffusion models (dMLLMs). We introduce a unified mathematical framework that clarifies their historical development and establishes a principled taxonomy. Our core paradigm centers on full-attention-driven, multi-token parallel denoising generation, integrating discrete probabilistic modeling, token-level noise scheduling, multi-stage training, and cross-modal alignment. The survey encompasses over 100 open-source and industrial models, demonstrating that dLLMs/dMLLMs match or approach autoregressive models’ performance across language, vision-language, and biological sequence tasks—while achieving up to 10× inference speedup, significantly enhanced output controllability, and improved dynamic response capability.
Existing discrete diffusion language models predominantly adopt full-decoder architectures, where each denoising step requires executing the entire network, resulting in high computational overhead and inefficient inference. Method: We propose the first encoder-decoder-based discrete diffusion model: a dedicated encoder learns clean-text representations, while a lightweight decoder performs iterative denoising; combined with block-wise sequence partitioning and specialized training/sampling algorithms, this design decouples representation learning from noise removal. Contribution/Results: Our architecture significantly improves training stability and inference throughput. Empirical evaluation on summarization, machine translation, and mathematical reasoning demonstrates superior quality–latency trade-offs at reduced computational cost. This work establishes a new paradigm for efficient discrete diffusion modeling.
This work addresses the inherent trade-offs among generation quality, diversity, and inference efficiency between autoregressive (AR) and diffusion-based sequence generation paradigms. To unify these frameworks, we propose position-specific noise hyperschedules that parameterize both AR and diffusion processes within a single formulation; design a hybrid token-level noising mechanism that dynamically balances absorbing-noise and uniform-noising strategies to enable error correction; and introduce KV-cache-adapted attention masking to accelerate parallel decoding. Experiments on standard language modeling benchmarks demonstrate state-of-the-art perplexity, along with significant improvements in generated sequence diversity, fidelity, and robustness—while simultaneously reducing inference latency.
Masked diffusion language models (MDMs) exhibit significant performance disparities across evaluation metrics, yet their fundamental trade-offs between efficiency and accuracy remain theoretically uncharacterized. Method: We establish the first rigorous theoretical framework for MDMs, integrating probabilistic modeling, information-theoretic bounds, and sampling complexity analysis to systematically characterize their intrinsic capability limits. Results: We prove that under perplexity, MDMs achieve near-optimal performance in a constant number of steps—independent of sequence length—whereas under sequence error rate, sampling steps must scale linearly with length. This reveals the critical insight that parallel sampling does not universally improve efficiency, challenging prevailing intuitions. All theoretical findings are empirically validated across diverse architectures and datasets. Our work provides both a foundational theory and practical guidance for the design, analysis, and evaluation of diffusion-based language models.
Autoregressive (AR) text generation in large language models (LLMs) suffers from slow inference due to sequential token prediction. Method: This paper systematically surveys and restructures the parallel text generation landscape, proposing the first unified taxonomy encompassing AR parallel decoding, non-autoregressive (Non-AR) modeling, diffusion-based language models, and knowledge distillation. Through theoretical analysis and empirical benchmarking across standard datasets and real-world scenarios, it characterizes the speed–quality–efficiency trade-offs inherent in each paradigm. Contribution/Results: The study identifies key acceleration pathways and synergistic integration opportunities, establishes performance boundaries, summarizes state-of-the-art advances, and highlights persistent challenges—including scalability, output consistency, and generalization. It delivers the first structured technical roadmap for efficient LLM inference, advancing the paradigm shift from serial to parallel generation.
This paper addresses the lack of convergence theory for diffusion language models (DLMs). Methodologically, it establishes the first rigorous asymptotic characterization of sampling error from an information-theoretic perspective, deriving tight upper and lower bounds on the sampling error in terms of KL divergence. It proves that the error decays at rate $1/T$ with respect to the number of iterations $T$, where the leading constant is proportional to the mutual information among tokens within the sequence. The bound is both tight and interpretable, revealing a fundamental trade-off between parallel generation efficiency and sequential dependency structure. Contributions include: (1) the first precise theoretical analysis of DLM sampling convergence; (2) the first incorporation of mutual information as a key complexity measure governing the convergence rate; and (3) the first rigorous theoretical foundation for efficient parallel text generation, thereby filling a critical theoretical gap in diffusion-based language modeling.
This work addresses the substantial computational redundancy in diffusion language models during inference, where full-sequence attention is repeatedly computed—even over already decoded or masked regions. The study is the first to reveal structural locality and temporal stability in the decoding process and introduces a training-free sliding window mechanism that dynamically partitions tokens into active, buffer, and far-field regions. Attention is computed only within a localized window, complemented by token-level pruning, KV cache reuse, and a phased refresh strategy. The method is directly applicable to pretrained models and achieves up to 99× inference speedup on LLaDA and Dream while largely preserving generation quality under the same computational budget.
This study addresses the lack of systematic evaluation of diffusion language models (DLMs) across diverse tasks, architectures, and inference configurations, which hinders their practical deployment. For the first time, it presents a comprehensive benchmark of eight state-of-the-art DLMs within a unified framework, evaluating their performance and computational efficiency across eight standard tasks under varying inference budgets. The analysis rigorously assesses generation quality and efficiency while dissecting the impact of key inference factors—such as denoising step count, context length, and parallel de-masking strategies—on model behavior. The findings demonstrate that inference-stage design critically governs the trade-off between performance and efficiency, clarifying the scenarios where DLMs excel or fall short, and offering actionable guidelines for optimizing their real-world application.
This study addresses the inefficiency of token-by-token generation in diffusion language models for long sequences by proposing the BLD framework. This method introduces a novel block-level denoising mechanism that compresses over a thousand consecutive tokens into a small number of block latent variables, combined with branched decoding to enable parallel local autoregressive generation. By integrating continuous diffusion, latent compression, and grouped conditional generation techniques, BLD effectively preserves both textual fluency and diversity while substantially accelerating long-text inference. Specifically, the proposed framework reduces computational costs by 80-fold and increases throughput by more than six times, offering a highly efficient solution for scaling diffusion-based language models to longer contexts.
Existing diffusion language models rely solely on local token information during sampling, neglecting global sequence structure and thus struggling to balance generation quality with parallel efficiency. This work formulates the sampling order selection as an NP-hard optimization problem for the first time and introduces Attn-Sampler, a training-free algorithm that leverages a computationally tractable approximation based on descending column sums of the attention matrix. By integrating attention mechanism analysis, sampling rank approximation, and dynamic thresholding for acceleration, the proposed method significantly outperforms baseline approaches such as greedy search across multiple benchmarks. It simultaneously enhances both text generation quality and parallelizability, offering a theoretically grounded and practically effective framework for attention-guided sampling in diffusion-based language models.
This work addresses the underutilization of low-confidence tokens discarded during the denoising process of discrete diffusion language models, which often leads to delayed and insufficient evidence retrieval in retrieval-augmented generation (RAG). To overcome this limitation, the authors propose SARDI—a dynamic RAG framework that repurposes these discarded tokens as proactive retrieval signals to guide external knowledge acquisition early in the generation process. SARDI requires no additional training and is compatible with any off-the-shelf retriever and inference-time discrete diffusion model. Experimental results demonstrate that SARDI consistently outperforms existing training-free diffusion-based and autoregressive RAG approaches across five challenging multi-hop question answering benchmarks, achieving up to an 8× improvement in throughput.