Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the diversity collapse in uniform-state diffusion language models, which stems from their inability to effectively preserve correct tokens during generation. Through an analysis grounded in NELBO decomposition, this work identifies "correct token retention" as a critical missing component and proposes CTR-Reg, a regularized auxiliary loss function. This approach guides the model to actively retain unperturbed correct tokens during denoising, achieving a balance between self-correction and generation diversity without modifying existing samplers. Experimental results demonstrate that CTR-Reg improves clean token accuracy by 26.5% on average, enables single-step revision to converge within three to eleven positions, and halves generation perplexity while maintaining high diversity.
📝 Abstract
Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.
Problem

Research questions and friction points this paper is trying to address.

Uniform-state diffusion models
Correct-token retention
Self-correction
Sample diversity
Diffusion language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uniform-state diffusion models
Correct-token retention
Self-correction
Regularization
Diffusion language models
🔎 Similar Papers