IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Masked discrete diffusion models are difficult to fine-tune with reward signals due to costly iterative sampling and intractable sequence likelihoods. To address this, we propose the IDRF framework, which replaces KL penalties with inverse distillation regularization and enables few-step generation fine-tuning by modeling a finite-horizon Markov decision process and optimizing trajectory proxy losses via truncated policy gradients. Theoretically, we prove that an upper bound on the inverse distillation loss constrains the KL divergence, preserving the student model’s few-step sampling capability without requiring reference model rollouts. Experiments demonstrate that our approach reduces denoising steps by 32× across DNA, image, and text generation tasks, achieving high rewards while effectively mitigating reward hacking and maintaining sample quality.
📝 Abstract
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
Problem

Research questions and friction points this paper is trying to address.

masked discrete diffusion models
reward fine-tuning
few-step generation
intractable sequence likelihoods
iterative sampling cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Discrete Diffusion Models
Inverse-Distillation Regularization
Reward Fine-tuning
Few-step Generation
Policy Gradient
🔎 Similar Papers