An End-to-End Latent-Rollout Approach for Pushing Few-Step ImageNet-$256$ Generation to FID $1.11$ without Fr\'echet Losses

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of iterative joint optimization and the computational complexity of explicit FrΓ©chet distance loss in few-step image generation by proposing the Flow-in-Stage Transformer (FiST) architecture. FiST adopts a "distill-then-refine" strategy to enable decoder-free end-to-end gradient propagation and full trajectory optimization within the latent space. By integrating a REPA-pretrained SiT backbone, cross-stage hidden communication, and adversarial with auxiliary classification objectives, the framework enhances generation quality using only distribution-level supervision. Evaluated on ImageNet-256, FiST achieves state-of-the-art performance with FID scores of 1.11 and 1.15 for three-step and two-step generation, respectively, demonstrating the viability of efficient few-step synthesis.
πŸ“ Abstract
Iterative generation poses a joint optimization problem across steps, as intermediate predictions shape subsequent computations and ultimately determine the final output distribution. Few-step generators distilled from pretrained diffusion and flow-matching models make such optimization computationally practical end to end. We build on this opportunity with a distill-then-refine approach that uses teacher imitation to establish a strong initialization for a few-step rollout in latent space, then shifts to end-to-end refinement of the complete latent rollout against real data. We introduce FiST (Flow-in-Stage Transformer), an architecture that composes learned latent-state transitions in a few stages using a shared Transformer, with optional cross-stage hidden communication. Distillation applies teacher-forced regression to selected states along teacher trajectories; refinement replaces this supervision with adversarial and auxiliary classification objectives on the final latent output. A trainable discriminator module operates on semantically rich features extracted from clean real and generated latents by a frozen SiT backbone pretrained with REPA. All training takes place in latent space, without image decoding. During refinement, FiST consumes its own intermediate predictions, and endpoint gradients pass through every generation stage. For class-conditional generation on ImageNet at $256\times256$, our approach achieves FID 1.11 (IS 282) with three stages and FID 1.15 (IS 280) with two. These results demonstrate competitive few-step generation through learned distribution-level supervision, without explicit Fr\'echet-distance minimization. Ablations characterize how distillation, pretrained checkpoint choices, refinement supervision, and cross-stage hidden communication affect generation quality.
Problem

Research questions and friction points this paper is trying to address.

Few-step generation
Iterative generation
End-to-end optimization
Latent rollout
Image synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent-Rollout
FiST
Few-Step Generation
Distill-then-Refine
End-to-End Optimization
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.