π€ AI Summary
This study addresses the challenges of complex trajectories and difficult schedule design when flow matching models leverage pretrained representations for collaborative denoising. We propose a novel approach that directly predicts and conditions on DINO representations. Guided by the principle that predictable targets require no denoising, our method eliminates dual denoising trajectories, secondary ODEs, and dependencies on specific noise schedules, thereby establishing a concise and efficient representation-guidance mechanism. On ImageNet, this approach achieves faster latent-space convergence than state-of-the-art methods while requiring only half the training epochs, and improves pixel-space FID by over 20%. These results demonstrate substantial enhancements in both training efficiency and generation quality.
π Abstract
Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules. We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction. This removes the need for a second ODE and any representation-specific denoising schedules, while retaining the benefits of representation guidance. Our approach converges substantially faster and achieves better generation quality as measured by FID score. On ImageNet, it outperforms the state of the art in latent space at 2x fewer epochs than prior methods; in pixel space, it improves FID over comparable prior methods by more than 20%. These results support a simple principle: do not denoise what you can predict. Our code is openly available at https://github.com/arijit-hub/dino_forcing.