🤖 AI Summary
This study addresses the scarcity of ideal samples and the high cost of conditional estimation in reward-guided diffusion model fine-tuning by proposing the DOHF algorithm. Grounded in the Doob h-transform, this method converts probabilistic conditioning into an online training strategy, achieving efficient alignment through optimal weight allocation and locally corrected distillation. A core contribution lies in providing a unified perspective that explains both DiffusionNFT and classifier-free guidance, while supporting black-box, non-differentiable rewards without requiring auxiliary evaluation networks. Experimental results demonstrate that DOHF significantly enhances the alignment performance of diffusion models across diverse scenarios, including statistical sampling and visual generation.
📝 Abstract
Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.