One-Step Generative Modeling via Unbalanced Optimal Transport

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of balanced optimal transport in mini-batch training, where enforced mass matching renders generative flows highly sensitive to data composition. To overcome this, we propose an unbalanced optimal transport gradient flow framework that integrates Vlasov-Fokker-Planck dynamics with Diffusion Transformer (DiT) architectures. The core insight is that generated and real samples should be treated asymmetrically; specifically, we achieve one-sided optimization by fixing the generated marginal while relaxing only the real data marginal. Both theoretical analysis and empirical evaluations demonstrate that this one-sided relaxation strategy yields superior robustness compared to two-sided relaxation and balanced transport formulations. Notably, our approach achieves a one-step generation FID of 1.22 on ImageNet-256, establishing a new state-of-the-art record in the field.
📝 Abstract
Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches of generated and real samples, which provide only imperfect approximations to the underlying distributions. Balanced optimal transport enforces exact mass matching within every mini-batch, making the estimated field sensitive to the particular composition of the real-data batch. We find that generated and real samples should be treated asymmetrically: letting the mass assigned to real samples adapt while keeping every generated sample fully transported improves generation across six feature-space metrics in controlled ablations, and is more robust to the relaxation strength than relaxing both marginals simultaneously, which falls below balanced transport under stronger relaxation. Motivated by this observation, we propose Unbalanced Optimal Transport Gradient Flow (UOT-GF), which keeps the generated-sample marginal fixed and relaxes only the real-data marginal. Under identical settings at DiT-B/2 on ImageNet-256, UOT-GF improves Fr\'echet Inception Distance (FID) from 1.53 to 1.46 over the balanced W-Flow baseline; scaling the same recipe yields 1.34 and 1.22 FID at L/2 and XL/2, the best FID among the one-step models we compare. We further derive the induced UOT transport force, establish a kinetic Vlasov--Fokker--Planck formulation whose overdamped zero-temperature limit recovers the drifting dynamics, and characterize non-target stationary states together with sufficient conditions for convergence.
Problem

Research questions and friction points this paper is trying to address.

one-step generative modeling
unbalanced optimal transport
drifting models
mini-batch approximation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unbalanced Optimal Transport
One-Step Generation
Gradient Flow
Diffusion Models
Vlasov-Fokker-Planck
🔎 Similar Papers