🤖 AI Summary
This study addresses the limitation that confidence sampling in diffusion language models degrades generation diversity and constrains reinforcement learning-based post-training. To overcome this, we propose a hybrid sampling strategy that introduces autoregressive sequential decoding exclusively at low-confidence steps while preserving parallel generation elsewhere. Furthermore, we present the ForkGRPO algorithm to optimize policy ratios and reduce post-training costs. Experimental results demonstrate that our approach matches the generation diversity of autoregressive models while achieving a 2–3× improvement in inference efficiency and substantially lowering training overhead. Overall, this work effectively balances generation quality, diversity, and computational efficiency for diffusion language models.
📝 Abstract
Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.