Scaling and Distilling Text Embeddings for Better Diffusibility

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that overly discriminative embedding spaces in continuous diffusion language models render the sampling process prone to failure. To mitigate this issue, we extend continuous diffusion language models based on the T5Gemma architecture and propose a decoding probability soft-label knowledge distillation method. By optimizing the connectivity of the embedding space, our approach constructs latent representations better suited for the diffusion process, effectively resolving the sampling difficulties inherent in diffusion-based generation. Empirically, the proposed medium-sized model achieves a generative perplexity of 17.8 on the OpenWebText benchmark, outperforming GPT-2-M. These results validate the effectiveness of soft-label distillation in enhancing the generative quality of diffusion language models.
📝 Abstract
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Language Models
Text Embeddings
Diffusibility
Latent Space
Continuous Diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Knowledge Distillation
Text Embeddings
Latent Space
Soft Labels
🔎 Similar Papers
No similar papers found.