🤖 AI Summary
This work addresses the slow convergence and degraded generation quality in standard Flow Matching training caused by gradient conflicts arising from data heterogeneity. It is the first to model the Flow Matching objective as a dynamic quadratic form dominated by the Neural Tangent Kernel (NTK), thereby revealing how heterogeneous data interact within the residual vector field. Building on this insight, the authors propose a semantic granularity alignment strategy that explicitly modulates feature cross-terms to mitigate gradient interference. The method significantly accelerates convergence on both DiT and U-Net architectures while enhancing the structural integrity of generated images, achieving a superior trade-off between training efficiency and sample quality.
📝 Abstract
In this work, we analyze the optimization dynamics of generative fine-tuning. We observe that under the Flow Matching framework, the standard MSE objective can be formulated as a Quadratic Form governed by a dynamically evolving Neural Tangent Kernel (NTK). This geometric perspective reveals a latent Data Interaction Matrix, where diagonal terms represent independent sample learning and off-diagonal terms encode residual correlation between heterogeneous features. Although standard training implicitly optimizes these cross-term interferences, it does so without explicit control; moreover, the prevailing data-homogeneity assumption may constrain the model's effective capacity. Motivated by this insight, we propose Semantic Granularity Alignment (SGA), using Text-to-Image synthesis as a testbed. SGA engineers targeted interventions in the vector residual field to mitigate gradient conflicts. Evaluations across DiT and U-Net architectures confirm that SGA advances the efficiency-quality trade-off by accelerating convergence and improving structural integrity.