Less Uniform Discrete Diffusion is More Powerful and Scalable

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of overly uniform training objectives and condition-target confusion during sampling when scaling uniform diffusion language models. To this end, it proposes the LUDI framework. Methodologically, a non-uniform loss function is designed to guide reverse transitions, while token-level time embeddings are introduced to optimize supervision signals. Additionally, a confidence-based few-step sampling mechanism is developed to enhance generation efficiency. Leveraging continuous training techniques, a 7B autoregressive model is converted into LUDI-7B, which integrates token-level corruption prompts for efficient discrete diffusion modeling. Experimental results demonstrate that this framework achieves a decoding acceleration of three tokens per step and yields performance comparable to masked diffusion baselines on complex reasoning tasks, thereby significantly improving the scalability and inference potential of discrete diffusion models.
📝 Abstract
Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.
Problem

Research questions and friction points this paper is trying to address.

Uniform Diffusion Language Models
Scaling
Training Objective
Condition-Target Confusion
Sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Less Uniform Diffusion
Discrete Diffusion Language Models
Per-token Time Embeddings
Few-step Sampling
Scalability
🔎 Similar Papers
2022-09-02ACM Computing SurveysCitations: 1628
Kaibo Wang
Kaibo Wang
Tsinghua University
Statistical Quality ControlIndustrial Engineering
D
Ding Ding
The Hong Kong University of Science and Technology, Huawei Foundation Model Department
F
Fangyu Ding
The Hong Kong University of Science and Technology, Huawei Foundation Model Department
Zijin Feng
Zijin Feng
The Chinese University of Hong Kong
Large Language ModelsData Mining
Han Shi
Han Shi
HKUST, Huawei Noah's Ark Lab
Machine LearningAI4MATH
H
Haili Bai
Huawei Foundation Model Department
J
Jiacheng Sun
Huawei Foundation Model Department
Y
Yang Xiang
The Hong Kong University of Science and Technology