Targeted Remasking: Replacing Token Editing with Token-to-Mask Refinement in Discrete Diffusion Language Models

📅 2026-04-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Discrete diffusion language models suffer from contextual contamination, entangled error detection and replacement, and a mismatch between training and inference noise due to direct token replacement during decoding. This work proposes Token-to-Mask (T2M), a plug-and-play remasking mechanism that resets suspected erroneous tokens to the mask token, enabling the model to re-predict within a cleaner context without requiring retraining. T2M innovatively decouples error detection from replacement and introduces three complementary token-level refinement strategies—probabilistic, trigger-mirroring, and temporal-difference—to transform systematic errors into mask-like noise familiar to the model. Evaluated across twelve benchmarks spanning knowledge, reasoning, mathematics, code generation, and instruction following, T2M demonstrates consistent gains, notably improving mathematical performance by +5.92% on CMATH and correcting 59.4% of end-of-sequence token errors.
📝 Abstract
Discrete masked diffusion language models such as LLaDA generate text through iterative denoising, where mask tokens are progressively replaced with predicted tokens. LLaDA2.1 introduced a Token-to-Token (T2T) editing mechanism that accelerates generation by directly replacing committed tokens suspected of being incorrect. However, we identify fundamental limitations of T2T editing: it couples error detection with replacement, pollutes the generation context with potentially incorrect tokens, and introduces a train-inference noise mismatch where systematic model-generated errors differ from the random perturbations seen during training. We propose Token-to-Mask (T2M) remasking, a training-free, drop-in replacement for T2T editing that resets suspected erroneous tokens back to the mask state, allowing the diffusion process to re-predict them under cleaner context. We design and empirically validate three complementary error detection strategies -- probability-based, trigger-mirrored, and temporal-difference-based -- and provide a unified theoretical analysis showing that T2M remasking purifies the generation context, converts systematic inference errors back to the model's native mask noise type, and enables delayed commitment for joint multi-position optimization. Comprehensive experiments across 12 benchmarks spanning knowledge, reasoning, mathematics, coding, and instruction following show that T2M generally improves performance on tasks requiring precise token-level output, with the largest gain on mathematics (+5.92% on CMATH). Error analysis on CMATH reveals that the dominant failure mode is last-mile token corruption -- where correct reasoning produces a corrupted final answer -- and that T2M repairs 59.4% of such cases.
Problem

Research questions and friction points this paper is trying to address.

discrete diffusion language models
token editing
train-inference mismatch
mask tokens
error correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-to-Mask
discrete diffusion language models
remasking
error correction
mask-based refinement