Discrete Stochastic Localization for Non-autoregressive Generation

📅 2026-02-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Non-autoregressive generation often suffers from low sampling efficiency due to error accumulation and distributional shift. This work proposes the Denoising with Shared Latents (DSL) method, which trains a single SNR-invariant denoiser within a Diffusion Transformer to jointly handle partially observed text across continuous noise levels, integrating intermediate draft noise with a masked endpoint corruption mechanism. Requiring only modifications to model training—without complex sampling strategies—DSL substantially improves step efficiency, self-correction capability, and uncertainty calibration. On OpenWebText, DSL achieves a higher MAUVE score than MDLM+ReMDM using only about one-quarter of the denoiser calls, and under high computational budgets, its generation quality matches that of autoregressive models.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Diffusion Models for Vision

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Non-autoregressive (NAR) generation reduces decoding latency by predicting many tokens in parallel, but iterative refinement often suffers from error accumulation and distribution shift under self-generated drafts. Masked diffusion language models (MDLMs) and their remasking samplers (e.g., ReMDM) can be viewed as modern NAR iterative refinement, where generation repeatedly revises a partially observed draft. In this work we show that \emph{training alone} can substantially improve the step-efficiency of MDLM/ReMDM sampling. We propose \textsc{DSL} (Discrete Stochastic Localization), which trains a single SNR-invariant denoiser across a continuum of corruption levels, bridging intermediate draft noise and mask-style endpoint corruption within one Diffusion Transformer. On OpenWebText, \textsc{DSL} fine-tuning yields large MAUVE gains at low step budgets, surpassing the MDLM+ReMDM baseline with \(\sim\)4$\times$ fewer denoiser evaluations, and matches autoregressive quality at high budgets. Analyses show improved self-correction and uncertainty calibration, making remasking markedly more compute-efficient.
Problem

Research questions and friction points this paper is trying to address.

Non-autoregressive generation
error accumulation
distribution shift
iterative refinement
masked diffusion language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Stochastic Localization
Non-autoregressive Generation
Masked Diffusion Language Models
SNR-invariant Denoiser
Iterative Refinement