Decoding in Order-Agnostic Language Models: Chain-Rule Deviation and Uniform Spreading

📅 2026-05-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that order-agnostic language models (OALMs) exhibit path-dependent artifacts in likelihood estimation across different revelation orders, obscuring true content difficulty. The authors propose the variance of confidence trajectories as a novel diagnostic metric for decoding path quality and theoretically demonstrate that, under fixed total likelihood, uniform per-step confidence maximizes target recoverability. Building upon the discrete diffusion language model (dLLM) framework and integrating confidence-first (CF) decoding with chain-rule bias analysis, they empirically validate on C4 and four downstream tasks that low confidence variance effectively identifies structured decoding paths and exhibits a strong positive correlation with task accuracy.
📝 Abstract
Order-agnostic language models (OALMs), including discrete diffusion language models (dLLMs), are trained to predict masked tokens under arbitrary conditioning sets, allowing sequences to be generated or scored under arbitrary reveal orders at inference time. In LLaDA-2.1, we report three findings. First, the learned conditionals are not exact factorizations of a coherent joint distribution: changing only the reveal order shifts target log-likelihood by up to 0.49 nats/token, so likelihood alone mixes content difficulty with path-dependent artifacts. Second, although confidence-first (CF) decoding is order-agnostic, its reveal orders are close to left-to-right (L2R) on content tokens. Third, we propose a complementary diagnostic based on the shape of the confidence trace. A uniform-spreading theorem shows that, at fixed total likelihood, target recoverability is maximized when per-step confidence is spread uniformly; the resulting deviation motivates $\mathrm{Var}(\log q_t)$ as a diagnostic for comparing decoding paths. Across C4 and four downstream benchmarks, low variance separates structured paths from random ordering, and variance is consistently associated with downstream correctness. These results support reporting mean confidence and confidence variance jointly when comparing OALM decoding paths.
Problem

Research questions and friction points this paper is trying to address.

order-agnostic language models
reveal order
likelihood inconsistency
decoding paths
confidence variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

order-agnostic language models
chain-rule deviation
uniform spreading
confidence variance
decoding paths
🔎 Similar Papers
No similar papers found.