Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of diffusion language model agents falling into retry loops during embodied tasks due to masked decoding. It reveals a task-agnostic bias mechanism induced by context saliency. To mitigate this, the work proposes a training-free reverse scoring method that eliminates bias factors through reverse conditional probability analysis, thereby correcting action distributions during parallel decoding. Experimental results demonstrate that this approach significantly improves both task success rates and progress rates across four multi-turn embodied benchmarks. Ultimately, this research provides an efficient, training-free error correction mechanism for the masked denoising process in diffusion language models.
📝 Abstract
Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.
Problem

Research questions and friction points this paper is trying to address.

diffusion language models
embodied agents
retry loops
masked decoding
reverse scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Reverse Scoring
Masked Decoding
Embodied Agents
Training-free Remedy
🔎 Similar Papers
No similar papers found.