Unifying Distributional Training for One-Step Visual Generation

📅 2026-09-28
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified theoretical framework for distribution matching and the issue of mode collapse in one-step visual generation. To this end, we propose MGFlow, a method grounded in Wasserstein gradient flows that integrates optimal transport with score matching theory to construct a unified feature distribution modeling framework. By introducing a mass-constrained allocation mechanism, MGFlow effectively mitigates mode collapse while supporting granularity adjustment via Gaussian mixture models. Experimental results demonstrate that MGFlow achieves state-of-the-art performance on ImageNet with an FDr of 1.45. Notably, its single-step generation quality surpasses that of the original four-step FLUX.2 model, providing both a rigorous theoretical foundation and a practical solution for efficient visual generation.
📝 Abstract
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/
Problem

Research questions and friction points this paper is trying to address.

one-step visual generation
distributional training
mode collapse
Wasserstein gradient flow
feature matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributional Training
One-Step Visual Generation
Wasserstein Gradient Flow
Gaussian Mixture Modeling
MGFlow
🔎 Similar Papers
No similar papers found.