Geometry-Aware Preference Optimization for Text-to-Image Diffusion Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing diffusion model preference alignment methods that neglect image manifold geometry, resulting in mismatched optimization dynamics and degraded generation quality and diversity. This work proposes Anisotropic Geometry-Aware Preference Optimization (APO), which reformulates Direct Preference Optimization (DPO) from a manifold geometric perspective for the first time. By leveraging the reference model to derive an adaptive anisotropic metric rather than relying on uniform Euclidean processing, APO precisely balances tangential semantic adjustments against hazardous normal-direction updates. Experimental results demonstrate that APO achieves an average win rate exceeding 60% across multiple benchmarks, significantly reduces the required training steps, and effectively preserves generation diversity throughout the entire training process.
📝 Abstract
Preference alignment has become a standard practice for text-to-image diffusion models. Direct Preference Optimization (DPO) simplifies this process by eliminating explicit reward modeling. Its diffusion variant, Diffusion-DPO, has become a widely adopted baseline. Diffusion-DPO essentially encourages the likelihood of preferred samples while suppressing dispreferred ones. In this paper, we revisit DPO-style alignment methods for diffusion models from the perspective of the manifold hypothesis. Under this view, natural images concentrate near a low-dimensional manifold embedded in the high-dimensional ambient space, whereas DPO directly optimizes preference distributions in the full space without accounting for this geometric structure. This creates a mismatch in the optimization dynamics: it suppresses geometry-preserving tangential updates, while insufficiently restricting hazardous normal-direction updates. This mismatch gradually degrades image quality and diversity. To address this issue, we propose Anisotropic Geometry-Aware Preference Optimization (APO), which replaces the uniform Euclidean treatment of prediction errors with a geometry-aware anisotropic metric derived from the reference model. Concretely, APO adaptively strengthens regularization in directions where the reference denoising function is highly sensitive, while relaxing constraints in directions that permit safe semantic adjustment. This recalibrates preference optimization according to the local manifold geometry, and maintains the original manifold structure. Experiments show that APO achieves strong performance and an average win rate exceeding 60\% against various existing alignment methods across diverse benchmarks. It requires significantly fewer training steps than prior methods, and preserves generation diversity throughout training.
Problem

Research questions and friction points this paper is trying to address.

Preference Optimization
Diffusion Models
Manifold Hypothesis
Text-to-Image Generation
Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference Optimization
Manifold Hypothesis
Anisotropic Metric
Diffusion Models
Geometry-Aware
🔎 Similar Papers
No similar papers found.