Representation-Space MMD for Diffusion Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low post-training efficiency and sampling trajectory dependency of diffusion language models by proposing a Maximum Mean Discrepancy (MMD) optimization framework operating within a frozen pretrained feature space. The method achieves distribution alignment by extracting multiple observations via a single forward pass, eliminating the need for full sampling or auxiliary models. Furthermore, it unifies discrete policy gradients with continuous direct differentiation and incorporates a hybrid-mask uniform diffusion mechanism to enable efficient optimization. Experimental results demonstrate that the proposed approach significantly reduces generation perplexity while improving both reasoning accuracy on GSM8K and the computation-accuracy trade-off. Notably, when applied to a 16B-parameter model, it effectively enhances decoding parallelism without compromising performance.
📝 Abstract
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Language Models
Maximum Mean Discrepancy
Post-training
Distribution Alignment
Decoding Parallelism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Maximum Mean Discrepancy
Representation Space
Post-training
Policy Gradients
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Ilya Drobyshevskiy
Yandex Research
I
Ilia Sudakov
Yandex Research
M
Maksim Semenov
HSE University
D
Denis Kuznedelev
Yandex Research
M
Maksim Ignatov
Yandex Research
P
Pavel Temirchev
Yandex Research
Nikita Balagansky
Nikita Balagansky
Central University
NLP
V
Viacheslav Meshchaninov
Constructor University
Nikita Gushchin
Nikita Gushchin
Skoltech
Generative ModelsDiffusion Models
Dmitry Baranchuk
Dmitry Baranchuk
Yandex Research
Generative ModelingComputer VisionSimilarity Search