🤖 AI Summary
This study addresses the intractability of class-wise coupling in conditional flow models under large-scale or continuous conditioning by proposing a global transport coupling method. Specifically, this work replaces conventional class-dependent coupling strategies with a class-agnostic global optimal transport algorithm and integrates classifier-free guidance (CFG) to optimize the training and inference of conditional generative models. The research reveals a critical phenomenon: while this coupling degrades performance in the absence of guidance, it yields substantial improvements when combined with CFG. Extensive experiments demonstrate that the proposed approach consistently enhances image generation quality across diverse data domains, model scales, and sampling budgets under both discrete categorical and continuous text conditioning.
📝 Abstract
Optimal-transport couplings have been shown to reduce training variance in unconditional flow models, but their role in conditional generation remains unclear. A natural approach constructs separate couplings for each condition, but this is impractical for large or continuous conditioning spaces found in modern image foundation models. We introduce Global Transport (GT), a global class-agnostic optimal-transport coupling, computed without class labels. GT can associate different conditions with different regions of the source noise, and consequently worsens performance without guidance. However, when combined with classifier-free guidance (CFG), GT consistently improves generation across domains, model scales, and sampling budgets. This reversal suggests that couplings for conditional flows should be evaluated both empirically and theoretically under the guided flow used at inference, rather than on unguided generation. We evaluate GT over both discrete class and continuous text conditioned image generation across model scales, and investigate how coupling choice alters guided trajectories. These results identify coupling design in the guided flow setting as a simple training time axis to improve performance without modifying existing architectures, samplers, or guidance mechanisms.