🤖 AI Summary
This study addresses the prediction unreliability in masked diffusion language models during parallel decoding, which stems from contextual uncertainty. To overcome this limitation, we propose a training-free reliable parallel decoding framework that transcends conventional fixed-block scheduling paradigms. Revealing that confidence alone is an insufficient evaluation criterion, our method innovatively incorporates inter-layer prediction stability analysis and cumulative entropy budget control mechanisms. By constraining upstream uncertainty to dynamically filter and commit candidate tokens, it achieves adaptive parallel decoding. Extensive evaluations on mathematical reasoning and code generation benchmarks demonstrate that the proposed approach attains optimal decoding throughput while preserving or even enhancing generation accuracy.
📝 Abstract
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.