Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of token inconsistency and unclear soft-token feedback mechanisms in parallel decoding for diffusion language models. By investigating soft-token feedback under frozen pretrained models, we identify a geometric mismatch between the feedback signals and the embedding space. Accordingly, we propose a training-free, geometry-aware embedding construction method that preserves non-principal components within the uncertainty to demonstrate that such feedback effectively enhances sequence coherence. Extensive experiments across four benchmarks show that our approach significantly outperforms standard parallel decoding and Euclidean baselines, substantially improving generation consistency. This work establishes a novel training-free optimization paradigm for efficient decoding in diffusion language models.
📝 Abstract
Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emph{soft tokens} can mitigate this issue by representing uncertain positions with continuous embeddings built from the model's predictive distribution at the previous decoding step. However, although soft tokens are commonly understood as preserving predictive uncertainty, how soft-token feedback improves parallel decoding has not been systematically examined. In this paper, we investigate this question in frozen pretrained DLMs to examine soft-token feedback without the effects of additional training. To construct soft-token inputs in a training-free setting, we identify a geometric mismatch between conventional soft-token construction and the pretrained embedding space. Based on this observation, we propose a training-free, geometry-aware construction of soft tokens. Our analysis of soft-token feedback suggests that uncertainty preservation alone does not fully explain how it reshapes subsequent predictions. To better explain how soft-token feedback improves parallel decoding, we provide empirical evidence that it favors coherent token sequences. Across four pretrained DLMs and four math and code benchmarks, our method outperforms standard parallel decoding and a training-free Euclidean soft-token baseline. Code: https://github.com/kodaikawamura/rethinking-soft-tokens
Problem

Research questions and friction points this paper is trying to address.

Diffusion Language Models
Parallel Decoding
Soft Tokens
Token Inconsistency
Geometric Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Soft Tokens
Parallel Decoding
Geometry-aware Construction
Training-free