diffusion language modeling

Designs, implements, and evaluates probabilistic sequence models and decoding algorithms that generate or recover natural-language token sequences by applying iterative diffusion/denoising processes instead of autoregressive factorization. This includes training objectives, architectures and sampling schedules for masked-token denoising, bidirectional/infill-capable decoding, and diffusion-based text generation, together with analyses of convergence, sampling speed, and tradeoffs versus autoregressive approaches.

diffusionlanguagemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Jun 16, 2025
RY
Runpeng Yu
🏛️ National University of Singapore

To address the limitations of autoregressive language models—including poor parallelizability, weak fine-grained controllability, and inadequate response awareness—this survey systematically reviews recent advances in discrete diffusion language models (dLLMs) and multimodal diffusion models (dMLLMs). We introduce a unified mathematical framework that clarifies their historical development and establishes a principled taxonomy. Our core paradigm centers on full-attention-driven, multi-token parallel denoising generation, integrating discrete probabilistic modeling, token-level noise scheduling, multi-stage training, and cross-modal alignment. The survey encompasses over 100 open-source and industrial models, demonstrating that dLLMs/dMLLMs match or approach autoregressive models’ performance across language, vision-language, and biological sequence tasks—while achieving up to 10× inference speedup, significantly enhanced output controllability, and improved dynamic response capability.

Analyze training, inference, and applications of diffusion modelsCompare discrete diffusion models with autoregressive modelsSurvey discrete diffusion models in language and multimodal tasks

Must-Read Papers

Most classic and influential ideas
View more

Encoder-Decoder Diffusion Language Models for Efficient Training and Inference

Oct 26, 2025
MA
Marianne Arriola
🏛️ Cornell University

Existing discrete diffusion language models predominantly adopt full-decoder architectures, where each denoising step requires executing the entire network, resulting in high computational overhead and inefficient inference. Method: We propose the first encoder-decoder-based discrete diffusion model: a dedicated encoder learns clean-text representations, while a lightweight decoder performs iterative denoising; combined with block-wise sequence partitioning and specialized training/sampling algorithms, this design decouples representation learning from noise removal. Contribution/Results: Our architecture significantly improves training stability and inference throughput. Empirical evaluation on summarization, machine translation, and mathematical reasoning demonstrates superior quality–latency trade-offs at reduced computational cost. This work establishes a new paradigm for efficient discrete diffusion modeling.

Accelerating discrete diffusion inference via encoder-decoder architectureImproving quality-throughput tradeoff in text generation tasksReducing computational cost in diffusion language model training

Unifying Autoregressive and Diffusion-Based Sequence Generation

Apr 08, 2025
NF
Nima Fathi
🏛️ ServiceNow Research

This work addresses the inherent trade-offs among generation quality, diversity, and inference efficiency between autoregressive (AR) and diffusion-based sequence generation paradigms. To unify these frameworks, we propose position-specific noise hyperschedules that parameterize both AR and diffusion processes within a single formulation; design a hybrid token-level noising mechanism that dynamically balances absorbing-noise and uniform-noising strategies to enable error correction; and introduce KV-cache-adapted attention masking to accelerate parallel decoding. Experiments on standard language modeling benchmarks demonstrate state-of-the-art perplexity, along with significant improvements in generated sequence diversity, fidelity, and robustness—while simultaneously reducing inference latency.

Introduce hyperschedules for distinct token noise schedulesPropose hybrid noising processes to fix past mistakesUnify autoregressive and diffusion models for sequence generation

Theoretical Benefit and Limitation of Diffusion Language Model

Feb 13, 2025
GF
Guhao Feng
🏛️ Peking University | Ant Group

Masked diffusion language models (MDMs) exhibit significant performance disparities across evaluation metrics, yet their fundamental trade-offs between efficiency and accuracy remain theoretically uncharacterized. Method: We establish the first rigorous theoretical framework for MDMs, integrating probabilistic modeling, information-theoretic bounds, and sampling complexity analysis to systematically characterize their intrinsic capability limits. Results: We prove that under perplexity, MDMs achieve near-optimal performance in a constant number of steps—independent of sequence length—whereas under sequence error rate, sampling steps must scale linearly with length. This reveals the critical insight that parallel sampling does not universally improve efficiency, challenging prevailing intuitions. All theoretical findings are empirically validated across diverse architectures and datasets. Our work provides both a foundational theory and practical guidance for the design, analysis, and evaluation of diffusion-based language models.

Analyze efficiency-accuracy trade-off in diffusion language models.Determine MDM's effectiveness and limitations compared to autoregressive models.Evaluate Masked Diffusion Model under different metrics.

A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models

Aug 12, 2025
LZ
Lingzhe Zhang
🏛️ Peking University | University of Illinois Chicago | Tsinghua University | XPENG | Alibaba Group | The Hong Kong University of Science and Technology (Guangzhou)

Autoregressive (AR) text generation in large language models (LLMs) suffers from slow inference due to sequential token prediction. Method: This paper systematically surveys and restructures the parallel text generation landscape, proposing the first unified taxonomy encompassing AR parallel decoding, non-autoregressive (Non-AR) modeling, diffusion-based language models, and knowledge distillation. Through theoretical analysis and empirical benchmarking across standard datasets and real-world scenarios, it characterizes the speed–quality–efficiency trade-offs inherent in each paradigm. Contribution/Results: The study identifies key acceleration pathways and synergistic integration opportunities, establishes performance boundaries, summarizes state-of-the-art advances, and highlights persistent challenges—including scalability, output consistency, and generalization. It delivers the first structured technical roadmap for efficient LLM inference, advancing the paradigm shift from serial to parallel generation.

Analyzing AR and Non-AR methods for speed-quality-efficiency trade-offsIdentifying challenges and future directions in parallel text generationSurveying parallel text generation techniques to overcome autoregressive bottlenecks

A Convergence Theory for Diffusion Language Models: An Information-Theoretic Perspective

May 27, 2025
GL
Gen Li
🏛️ The Chinese University of Hong Kong | University of Michigan

This paper addresses the lack of convergence theory for diffusion language models (DLMs). Methodologically, it establishes the first rigorous asymptotic characterization of sampling error from an information-theoretic perspective, deriving tight upper and lower bounds on the sampling error in terms of KL divergence. It proves that the error decays at rate $1/T$ with respect to the number of iterations $T$, where the leading constant is proportional to the mutual information among tokens within the sequence. The bound is both tight and interpretable, revealing a fundamental trade-off between parallel generation efficiency and sequential dependency structure. Contributions include: (1) the first precise theoretical analysis of DLM sampling convergence; (2) the first incorporation of mutual information as a key complexity measure governing the convergence rate; and (3) the first rigorous theoretical foundation for efficient parallel text generation, thereby filling a critical theoretical gap in diffusion-based language modeling.

Analyzes sampling error via KL divergence and mutual informationDevelops convergence guarantees for diffusion language modelsEstablishes tight upper and lower bounds for convergence

Latest Papers

What's happening recently
View more

This work addresses the substantial computational redundancy in diffusion language models during inference, where full-sequence attention is repeatedly computed—even over already decoded or masked regions. The study is the first to reveal structural locality and temporal stability in the decoding process and introduces a training-free sliding window mechanism that dynamically partitions tokens into active, buffer, and far-field regions. Attention is computed only within a localized window, complemented by token-level pruning, KV cache reuse, and a phased refresh strategy. The method is directly applicable to pretrained models and achieves up to 99× inference speedup on LLaDA and Dream while largely preserving generation quality under the same computational budget.

diffusion language modelsfull-sequence attentioninference acceleration

This study addresses the lack of systematic evaluation of diffusion language models (DLMs) across diverse tasks, architectures, and inference configurations, which hinders their practical deployment. For the first time, it presents a comprehensive benchmark of eight state-of-the-art DLMs within a unified framework, evaluating their performance and computational efficiency across eight standard tasks under varying inference budgets. The analysis rigorously assesses generation quality and efficiency while dissecting the impact of key inference factors—such as denoising step count, context length, and parallel de-masking strategies—on model behavior. The findings demonstrate that inference-stage design critically governs the trade-off between performance and efficiency, clarifying the scenarios where DLMs excel or fall short, and offering actionable guidelines for optimizing their real-world application.

computational efficiencyDiffusion Language Modelsevaluation protocols

This study addresses the inefficiency of token-by-token generation in diffusion language models for long sequences by proposing the BLD framework. This method introduces a novel block-level denoising mechanism that compresses over a thousand consecutive tokens into a small number of block latent variables, combined with branched decoding to enable parallel local autoregressive generation. By integrating continuous diffusion, latent compression, and grouped conditional generation techniques, BLD effectively preserves both textual fluency and diversity while substantially accelerating long-text inference. Specifically, the proposed framework reduces computational costs by 80-fold and increases throughput by more than six times, offering a highly efficient solution for scaling diffusion-based language models to longer contexts.

Diffusion Language ModelsEfficiencyLatent Compression

Existing diffusion language models rely solely on local token information during sampling, neglecting global sequence structure and thus struggling to balance generation quality with parallel efficiency. This work formulates the sampling order selection as an NP-hard optimization problem for the first time and introduces Attn-Sampler, a training-free algorithm that leverages a computationally tractable approximation based on descending column sums of the attention matrix. By integrating attention mechanism analysis, sampling rank approximation, and dynamic thresholding for acceleration, the proposed method significantly outperforms baseline approaches such as greedy search across multiple benchmarks. It simultaneously enhances both text generation quality and parallelizability, offering a theoretically grounded and practically effective framework for attention-guided sampling in diffusion-based language models.

diffusion language modelsgeneration qualityglobal sequence structure

This work addresses the underutilization of low-confidence tokens discarded during the denoising process of discrete diffusion language models, which often leads to delayed and insufficient evidence retrieval in retrieval-augmented generation (RAG). To overcome this limitation, the authors propose SARDI—a dynamic RAG framework that repurposes these discarded tokens as proactive retrieval signals to guide external knowledge acquisition early in the generation process. SARDI requires no additional training and is compatible with any off-the-shelf retriever and inference-time discrete diffusion model. Experimental results demonstrate that SARDI consistently outperforms existing training-free diffusion-based and autoregressive RAG approaches across five challenging multi-hop question answering benchmarks, achieving up to an 8× improvement in throughput.

diffusion language modelsdiscrete diffusionlookahead tokens

Hot Scholars

AF

Alexander Fraser

Technical University of Munich, Munich Center for Machine Learning, Munich Data Science Institute
Machine TranslationNatural Language ProcessingMachine LearningInformation Retrieval
GN

Goran Nenadic

Department of Computer Science, University of Manchester
Natural language processingtext mininghealth informatics
WG

Wenbo Guo

UC Santa Barbara
Machine LearningSecurity
QG

Quanquan Gu

Associate Professor of Computer Science, UCLA
AGILarge Language ModelsReinforcement LearningNonconvex Optimization
AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation