🤖 AI Summary
This study addresses velocity scaling in flow matching, which is known to enhance generative quality but has been misattributed in prior work to MSE-induced velocity underestimation. We provide the first rigorous characterization of its underlying mechanism, revealing that it fundamentally reduces population time lag, and establish a theoretical connection between this lag and optimal gain. Building upon these insights, we propose an adaptive gain selection algorithm based on lag measurement to optimize the sampling process. Extensive experiments demonstrate the effectiveness of our approach: under classifier-free guidance-free settings on ImageNet-256 with 25 neural function evaluations (NFEs), the FID score improves significantly from 28.0 to 12.2. These results validate both our theoretical framework and the practical utility of the proposed method for high-quality image generation.
📝 Abstract
Scaling a learned flow-matching velocity field $v_θ$ by a gain $γ(t)$ was recently shown to greatly improve generation quality. Prior work argued that velocity fields trained with mean-squared error (MSE) systematically underestimate velocity magnitude and that scaling corrects this error. We show that MSE training does not create a velocity-magnitude deficit. We find instead that velocity scaling reduces population time lag: sampled states at model time $t$ resemble training states from an earlier time. Velocity scaling and moving model time back are two ways to address this population time lag. Across architectures and model sizes, measuring population time lag and using it to select a gain greatly improves generation quality, reducing FID from 28.0 to 12.2 (estimated by linear interpolation between FID measurements at neighboring gains) on ImageNet-256 at NFE 25 without guidance.