Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of dynamic adaptability in inference strategies for masked diffusion language models. We propose a five-axis inference analysis framework based on strategy inversion, which precisely defines "adaptation opportunities," and design a lightweight, validation-set-calibrated detector for high-opportunity states to enable selective state adaptation. Experimental results demonstrate that, on the JSON infilling task using LLaDA-8B, adapting merely 10% of the states captures 56.9% of the optimal performance gain, significantly outperforming uniformly applied strategies. This work provides an efficient new paradigm for enhancing the generation efficiency of masked diffusion models.
📝 Abstract
Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes--score, cardinality, region, commitment, and planning--and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
Problem

Research questions and friction points this paper is trying to address.

Masked diffusion language models
state adaptation
unmasking strategy
strategy reversals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Diffusion Language Models
Selective Adaptation
Strategy Reversal
Adaptation Opportunity
Unmasking Strategy
🔎 Similar Papers
2024-01-31Workshop on Representation Learning for NLPCitations: 3