iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the degradation in generation diversity and quality caused by reward optimization during reinforcement learning post-training of diffusion models. To mitigate this issue, we propose a training method grounded in the incremental Feynman-Kac formula. Theoretically, this work corrects the prevailing assumption that updating only late timesteps is detrimental, establishing a novel framework for optimizing denoising diffusion policies that provides a rigorous theoretical foundation for the optimal trade-off between alignment and diversity. Empirically, the proposed approach significantly enhances both model alignment and generation diversity across three distinct tasks, thereby validating its effectiveness.
πŸ“ Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Policy Optimization
Alignment
Diversity
Reinforcement Learning
Reward Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Policy Optimization
Incremental Feynman-Kac Training
Alignment-Diversity Tradeoff
Reinforcement Learning
Timestep Analysis
πŸ”Ž Similar Papers
2024-07-16arXiv.orgCitations: 2