Embedding Prediction Helps Image Generation

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation in existing diffusion Transformers (DiTs) where conditional embeddings are statically reused and cannot adapt to evolving denoising states. To overcome this, we propose NEPA, which introduces a pioneering multi-embedding prediction mechanism and an embedding-conditional generation paradigm. Specifically, a Next-Embedding Predictive Autoregression Transformer dynamically generates conditioning signals, enabling the model to evolve its conditional inputs at each denoising step based on the current noise state, thereby transcending static conditioning constraints. When integrated with REPA, the resulting NEPA-DiT-XL achieves a FrΓ©chet Inception Distance (FID) of 1.32 on ImageNet while requiring only approximately one-third of the training compute compared to standard REPA.
πŸ“ Abstract
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
Image Generation
Conditional Embedding
Denoising Steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Next-Embedding Predictive Autoregression
Diffusion Transformers
Multi-Embedding Prediction
Embedding Conditioned Generation
Image Generation
πŸ”Ž Similar Papers
No similar papers found.