DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow convergence of diffusion models caused by high-compression-ratio image encoders and the inherent difficulty in balancing reconstruction fidelity with generation efficiency. To overcome these challenges, this work proposes a Decoupled Compact Semantic Autoencoder featuring a novel macro-micro decoupled architecture. Specifically, it integrates macro-level representations from a pretrained semantic encoder with micro-level details preserved by a pixel-level module, and is subsequently coupled with a Diffusion Transformer (DiT) for image generation. This approach effectively resolves the trade-off bottleneck between reconstruction quality and training speed under high compression ratios. Experimental results on ImageNet demonstrate superior performance, achieving a PSNR of 29.79 and a gFID of 3.37. These outcomes significantly outperform existing baselines while substantially accelerating the convergence of diffusion model training.
📝 Abstract
High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a $1.6$B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at $1024\times1024$ resolution.
Problem

Research questions and friction points this paper is trying to address.

high-compression tokenizer
diffusion model convergence
latent image generation
reconstruction fidelity
semantic autoencoder
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Autoencoder
High Compression
Diffusion Convergence
Decoupled Architecture
Image Generation