Alpha Diffusion Language Models: Factorization Alone Is Not the Problem

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prediction inconsistency in discrete diffusion models during parallel generation caused by reduced denoising steps, where conventional cross-entropy training struggles to ensure joint distribution consistency. We propose AlphaDLM, a framework that introduces a tunable sequence-level Alpha loss to optimize the joint distribution, theoretically analyzing how objective functions and factorization influence the fitted distribution to transcend cross-entropy limitations. Experiments based on the SDAR-1.7B architecture on TinyGSM demonstrate that our method preserves multiple valid completions while eliminating invalid combinations. With only four evaluation steps, AlphaDLM achieves 34.6% accuracy on GSM8K, significantly improving the trade-off between accuracy and computational efficiency in code and mathematical reasoning tasks.
📝 Abstract
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We identify conditions under which intermediate alpha preserves multiple valid completions while excluding invalid token combinations. Trained on TinyGSM, our method achieves 34.6% accuracy on GSM8K with only four model evaluations. We further scale the method to SDAR-1.7B and evaluate it on code and mathematics benchmarks. These results show that changing the training objective can improve the accuracy-computation trade-off of factorized diffusion language models.
Problem

Research questions and friction points this paper is trying to address.

discrete diffusion language models
parallel generation
joint prediction consistency
denoising steps
cross-entropy training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Alpha Diffusion Language Models
sequence-level alpha loss
discrete diffusion
parallel generation
factorization
🔎 Similar Papers
No similar papers found.