π€ AI Summary
This study addresses the poor generation quality of pixel-space flow models, attributing it for the first time to deficient discriminative geometry in their representations. To this end, we propose a persistent representation learning mechanism that continuously optimizes the discriminative geometric structure as the generator evolves, and theoretically establish the gradient equivalence between the kernel density estimation ratio loss and flow regression. This approach enables high-quality one-step generation without requiring a pretrained encoder. Experiments demonstrate that our method reduces the FrΓ©chet Inception Distance (FID) by 82%β95% compared to the original pixel-space flow model across multiple datasets. Furthermore, incorporating pretrained representations and velocity clipping yields additional performance improvements.
π Abstract
Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.