🤖 AI Summary
This study addresses the unresolved question of whether computational resources during drift model training should be prioritized for fitting the current target or for recomputing the drift field. We compare multi-step optimization with a fixed target against real-time drift field recomputation strategies on ImageNet. Our analysis reveals that resampling alters the target direction far more substantially than parameter updates do, motivating a training strategy that prioritizes frequent drift field refreshment under matched wall-clock time constraints. Experiments demonstrate that, given equivalent computational budgets, frequently recomputing the drift field significantly reduces FID compared to deeply optimizing a fixed target, thereby effectively improving both the generation quality and training efficiency of diffusion models.
📝 Abstract
Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.