🤖 AI Summary
Discrete diffusion models struggle to achieve efficient few-step sampling for language generation, limiting their practical deployment. This work proposes the Discrete MeanFlow Generator, extending MeanFlow to continuous-time Markov chains (CTMCs) for the first time. By defining a time-interval-averaged velocity field and deriving a self-consistency identity, we construct a theoretical framework based on transition kernel increments that yields training objectives compatible with standard paradigms while remaining computationally efficient. Experimental results demonstrate that this approach reduces the total variation distance by 67% in Potts simulations, achieves a 16× speedup on OpenWebText while maintaining the lowest perplexity, and matches state-of-the-art performance on ImageNet.
📝 Abstract
Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few-step generation remains a fundamental challenge. In this work, we introduce the Discrete Average Generator, a principled extension of MeanFlow to Continuous-Time Markov Chains (CTMCs). Analogously to how MeanFlow defines an average velocity field over a time interval in continuous spaces, we define an average generator as the normalized increment of the transition kernel over a time interval. We show that this average generator satisfies a self-consistency identity, which provides the foundation for our training objective. We further develop training strategies that align with the standard training paradigm of diffusion language models while keeping the resulting objective tractable. When projected onto per-coordinate marginals, the self-consistency identity admits a closed-form expression, enabling efficient training and inference. In Potts model simulations, our objective reduces the total variation distance of the $K$-step sampler by up to 67%. On OpenWebText, our method achieves the lowest generative perplexity among the evaluated methods for 8 to 64 sampling steps while enabling a $16\times$ acceleration, and achieves comparable performance to existing methods on ImageNet.