🤖 AI Summary
This study addresses the representation mismatch between reconstruction and denoising dynamics in two-stage latent diffusion models (LDMs). By revealing that LDMs are essentially autoencoders, this work proposes an end-to-end single-stage training framework. Methodologically, the DiT backbone is decomposed into mutually inverse encoding and decoding components, with full-timestep image-space supervision applied to align internal representations. This enables native joint learning in the latent space, entirely eliminating the need for independent pretraining. The proposed framework achieves FID scores of 1.80 and 1.90 on 256×256 and 512×512 class-conditional generation tasks, respectively, validating the effectiveness of the single-stage architecture.
📝 Abstract
Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.