🤖 AI Summary
This paper addresses the lack of convergence theory for diffusion language models (DLMs). Methodologically, it establishes the first rigorous asymptotic characterization of sampling error from an information-theoretic perspective, deriving tight upper and lower bounds on the sampling error in terms of KL divergence. It proves that the error decays at rate $1/T$ with respect to the number of iterations $T$, where the leading constant is proportional to the mutual information among tokens within the sequence. The bound is both tight and interpretable, revealing a fundamental trade-off between parallel generation efficiency and sequential dependency structure. Contributions include: (1) the first precise theoretical analysis of DLM sampling convergence; (2) the first incorporation of mutual information as a key complexity measure governing the convergence rate; and (3) the first rigorous theoretical foundation for efficient parallel text generation, thereby filling a critical theoretical gap in diffusion-based language modeling.
📝 Abstract
Diffusion models have emerged as a powerful paradigm for modern generative modeling, demonstrating strong potential for large language models (LLMs). Unlike conventional autoregressive (AR) models that generate tokens sequentially, diffusion models enable parallel token sampling, leading to faster generation and eliminating left-to-right generation constraints. Despite their empirical success, the theoretical understanding of diffusion model approaches remains underdeveloped. In this work, we develop convergence guarantees for diffusion language models from an information-theoretic perspective. Our analysis demonstrates that the sampling error, measured by the Kullback-Leibler (KL) divergence, decays inversely with the number of iterations $T$ and scales linearly with the mutual information between tokens in the target text sequence. In particular, we establish matching upper and lower bounds, up to some constant factor, to demonstrate the tightness of our convergence analysis. These results offer novel theoretical insights into the practical effectiveness of diffusion language models.