🤖 AI Summary
This study addresses the challenge of aligning few-step generative models via Direct Preference Optimization (DPO), where likelihood evaluation is infeasible due to their implicit nature. To overcome this bottleneck, we propose FestDPO, a framework that leverages non-parametric likelihood estimation and rapid sampling capabilities to construct a sample-level DPO loss approximation independent of model architectures and sampling procedures. Experimental evaluations demonstrate that FestDPO surpasses baseline methods in both win rate and human ratings for text-to-image generation. Furthermore, in protein backbone generation, it significantly improves β-sheet proportions and structural designability. These results validate the generalizability and effectiveness of the proposed approach across diverse domains, establishing a principled pathway for preference alignment in few-step generative modeling.
📝 Abstract
Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $β$-sheet fraction and better structural designability than the baselines.