🤖 AI Summary
This study addresses the prediction inconsistency in discrete diffusion models during parallel generation caused by reduced denoising steps, where conventional cross-entropy training struggles to ensure joint distribution consistency. We propose AlphaDLM, a framework that introduces a tunable sequence-level Alpha loss to optimize the joint distribution, theoretically analyzing how objective functions and factorization influence the fitted distribution to transcend cross-entropy limitations. Experiments based on the SDAR-1.7B architecture on TinyGSM demonstrate that our method preserves multiple valid completions while eliminating invalid combinations. With only four evaluation steps, AlphaDLM achieves 34.6% accuracy on GSM8K, significantly improving the trade-off between accuracy and computational efficiency in code and mathematical reasoning tasks.
📝 Abstract
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We identify conditions under which intermediate alpha preserves multiple valid completions while excluding invalid token combinations. Trained on TinyGSM, our method achieves 34.6% accuracy on GSM8K with only four model evaluations. We further scale the method to SDAR-1.7B and evaluate it on code and mathematics benchmarks. These results show that changing the training objective can improve the accuracy-computation trade-off of factorized diffusion language models.