Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study reveals that the Low-Confidence Remasking (LCR) mechanism in diffusion large language models severely compromises generation diversity. We demonstrate that this degradation stems from the over-filtering behavior of LCR rather than the generation order itself. To address this limitation, we propose a Top-Probability Position (TPP) decoding strategy coupled with an entropy-guided initialization method, thereby restoring the flexibility of arbitrary-order generation. Experimental results indicate that our approach achieves Pass@k performance comparable to autoregressive decoding while substantially enhancing rollout diversity, solution coverage, and downstream policy optimization outcomes. By decoupling generation diversity from rigid remasking heuristics, this work establishes a new paradigm for efficient and diverse generation in diffusion language models.
📝 Abstract
Masked diffusion language models support arbitrary-order generation, suggesting a natural way to produce diverse outputs. However, recent work argues that this flexibility reduces diversity by delaying high-uncertainty tokens that can lead to different generation paths. We trace this diversity loss not to arbitrary-order generation itself, but largely to low-confidence remasking (LCR), a widely used decoding rule. At each step, LCR samples a token at every masked position but commits only the sampled token with the highest probability, filtering out the rest. We show that this mechanism can exponentially suppress lower-probability tokens as more positions compete, and observe the same suppression in LLaDA. In contrast, top-probability position selection (TPP), which has often been conflated with LCR under the shared label confidence-based decoding, avoids this diversity loss. TPP first selects the position whose most likely token has the highest probability, then samples directly from that position's distribution. Replacing LCR with TPP restores diversity and yields Pass@$k$ comparable to left-to-right decoding, suggesting that the reported diversity loss stems largely from LCR's filtering rather than from generating high-confidence positions first. To further exploit order flexibility, we introduce Entropy-Guided Initialization (EGI), which samples the first token at the highest-entropy position and then follows TPP. This simple modification further improves rollout diversity and solution coverage beyond left-to-right decoding, with gains extending to downstream policy optimization, highlighting the potential of arbitrary-order generation for diverse rollouts.
Problem

Research questions and friction points this paper is trying to address.

Masked diffusion language models
generation diversity
arbitrary-order generation
low-confidence remasking
decoding strategies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Diffusion Language Models
Low-Confidence Remasking
Top-Probability Position Selection
Entropy-Guided Initialization
Arbitrary-Order Generation
🔎 Similar Papers
No similar papers found.