🤖 AI Summary
Masked discrete diffusion models are difficult to fine-tune with reward signals due to costly iterative sampling and intractable sequence likelihoods. To address this, we propose the IDRF framework, which replaces KL penalties with inverse distillation regularization and enables few-step generation fine-tuning by modeling a finite-horizon Markov decision process and optimizing trajectory proxy losses via truncated policy gradients. Theoretically, we prove that an upper bound on the inverse distillation loss constrains the KL divergence, preserving the student model’s few-step sampling capability without requiring reference model rollouts. Experiments demonstrate that our approach reduces denoising steps by 32× across DNA, image, and text generation tasks, achieving high rewards while effectively mitigating reward hacking and maintaining sample quality.
📝 Abstract
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.