🤖 AI Summary
This study addresses the miscalibration of softmax confidence during decoding in masked diffusion language models by proposing BayesER, a framework that leverages Bayesian predictive entropy to guide token commitment. Specifically, it constructs a lightweight posterior via training-free LoRA adapters and Laplace approximation, thereby optimizing position ranking and selection within denoising steps to enable uncertainty-aware decoding with cross-dataset transferability. Experimental results demonstrate that BayesER significantly reduces sequence-level calibration error while maintaining or improving accuracy on tasks such as code generation. Overall, this work establishes a reliable Bayesian inference decoding paradigm for diffusion language models.
📝 Abstract
Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be miscalibrated. We introduce BayesER (BAYESian Entropy-based Reordering), a post-hoc Bayesian decoding framework that uses predictive uncertainty to guide token commitment. In BayesER, we construct a lightweight approximate posterior centered at the pretrained checkpoint, similar to Laplace-LoRA but without training LoRA adapters. We average predictions over posterior samples and use predictive entropy to prioritize reliable positions. We examine how posterior predictions affect position ordering and token selection across benchmarks spanning code generation, mathematical reasoning, planning, and molecular generation. We show that BayesER reduces sequence-level calibration error while preserving or improving accuracy relative to common decoding schemes, including confidence-threshold decoding. Additionally, a posterior fitted on one code-generation dataset reduces calibration error on another without refitting, suggesting that Bayesian uncertainty may provide a transferable signal for more reliable MDLM decoding.