LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the representation mismatch between reconstruction and denoising dynamics in two-stage latent diffusion models (LDMs). By revealing that LDMs are essentially autoencoders, this work proposes an end-to-end single-stage training framework. Methodologically, the DiT backbone is decomposed into mutually inverse encoding and decoding components, with full-timestep image-space supervision applied to align internal representations. This enables native joint learning in the latent space, entirely eliminating the need for independent pretraining. The proposed framework achieves FID scores of 1.80 and 1.90 on 256×256 and 512×512 class-conditional generation tasks, respectively, validating the effectiveness of the single-stage architecture.
📝 Abstract
Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.
Problem

Research questions and friction points this paper is trying to address.

Latent Diffusion Models
Representation Mismatch
End-to-End Image Generation
Auto-Encoder
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Diffusion Model
Auto-Encoder
End-to-End Training
DiT Backbone Splitting
Diffusion-native Latent Space
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhengqiang Zhang
The Hong Kong Polytechnic University, OPPO Research Institute
Lingchen Sun
Lingchen Sun
The Hong Kong Polytechnic University
Computer VisionImage Processing
Rongyuan Wu
Rongyuan Wu
The Hong Kong Polytechnic University
Computational PhotographyGenerative Models
Q
Qiaosi Yi
The Hong Kong Polytechnic University, OPPO Research Institute
Xiangtao Kong
Xiangtao Kong
The Hong Kong Polytechnic University
image restoration
Chaodong Xiao
Chaodong Xiao
The Hong Kong Polytechnic University
Computer Vision
L
Lei Zhang
The Hong Kong Polytechnic University, OPPO Research Institute