Simplex Diffusion Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the information collapse in discrete diffusion models caused by hard sampling at intermediate steps, which impedes effective uncertainty propagation. To overcome this limitation, this work proposes Simplex Diffusion, which elevates the state space to a probability simplex to represent categorical belief distributions. By deriving closed-form reverse transitions optimized via cross-entropy loss and introducing a DDIM-like sampler with tunable stochasticity, the model preserves distributions rather than performing hard sampling during denoising, thereby enabling cross-step uncertainty propagation. Experimental results demonstrate that Simplex Diffusion achieves a GenPPL of 17.0 on OpenWebText and 49% accuracy in code generation. Furthermore, when distilled to eight steps, it attains a 32.1% problem-solving rate on GSM8K, significantly outperforming both baselines and comparative models employing substantially more sampling steps.
📝 Abstract
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).
Problem

Research questions and friction points this paper is trying to address.

Discrete Diffusion Models
Information Collapse
Categorical Sampling
Uncertainty Propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simplex Diffusion Models
Probability Simplex
Information Collapse
DDIM-like Sampler
Discrete Diffusion
🔎 Similar Papers