Remask, Don't Replace: Token-to-Mask Refinement in Masked Diffusion Language Models

📅 2026-04-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing masked diffusion language models suffer from three key limitations in token-to-token editing: difficulty in triggering corrections, interference from erroneous contextual tokens, and a mismatch between perturbation distributions during training and inference. This work proposes Token-to-Mask (T2M), a training-free and parameter-free remasking mechanism that resets suspicious token positions to the mask token, enabling their re-prediction in subsequent denoising steps based on more reliable context. T2M introduces, for the first time, a training-agnostic remasking strategy that replaces erroneous tokens with masks as conditioning signals—a formulation better aligned with semantic coherence in real-world inference—and provides theoretical justification for this approach. Integrated with three error-detection heuristics, T2M significantly improves token-level editing accuracy across eight benchmarks, yielding a 5.92-point gain on the CMATH task and effectively correcting 41.3% of “last-mile” errors.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Diffusion Models for VisionNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Masked diffusion language models such as LLaDA2.1 rely on Token-to-Token (T2T) editing to correct their own generation errors: whenever a different token crosses a confidence threshold, the committed token is overwritten. We identify three structural failure modes of this rule. The trigger cannot fire when no single alternative is confident enough; the replacement is computed under a context that may itself contain errors; and the uniform perturbations used to train the T2T stream do not resemble the coherent, semantically plausible mistakes that the model actually makes at inference. As an alternative, we propose Token-to-Mask (T2M) remasking. Rather than overwriting a suspect token with a new guess, T2M resets the position to the mask state, so that the next denoising step re-predicts it from an in-distribution context. The method is training-free, modifies only the editing rule, and introduces no new parameters. We pair it with three detection heuristics and give a short theoretical account of why a mask is a better conditioning signal than an erroneous token. Across 8 benchmarks, T2M improves accuracy on tasks that require exact token-level output. Its largest gain is +5.92 points on CMATH, where we attribute 79.9% of baseline errors to last-mile corruption (correct reasoning followed by a garbled final answer); T2M repairs 41.3% of these cases.
Problem

Research questions and friction points this paper is trying to address.

masked diffusion language models
Token-to-Token editing
generation errors
error correction
last-mile corruption
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-to-Mask
masked diffusion language models
error correction
training-free editing
in-distribution context