🤖 AI Summary
This study addresses the challenge of optimizing semantic correction and redrafting capabilities in discrete diffusion models, which is hindered by fixed or stochastic corruption processes. We propose the VSDD framework, formulating training as a Stackelberg game wherein the denoiser acts as the follower and the corruption strategy as the leader, thereby learning a semantics-aware corruption process driven by performance gains rather than reconstruction difficulty. By integrating variational inference with score function estimation, our method efficiently solves for Markovian corruption processes via single-step gradient approximation. Experimental results demonstrate that the proposed framework substantially improves molecular generation validity, reduces text perplexity, and enhances offline recommendation metrics.
📝 Abstract
Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.