🤖 AI Summary
Existing latent diffusion models (LDMs) suffer from a fundamental trade-off among latent-space smoothness, perceptual compression quality, and reconstruction fidelity—largely attributable to suboptimal autoencoder design. Method: We identify this architectural limitation and propose LDMAEs, a novel framework integrating Variational Masked Autoencoders (VMAEs) into the LDM paradigm. LDMAEs are the first to incorporate hierarchical feature modeling—inspired by Masked Autoencoders (MAEs)—into a variational autoencoding structure, enabling joint optimization of smoothness, perceptual quality, and fidelity directly in the latent space. VMAEs employ layered masking and reconstruction to learn compact, perception-driven representations, thereby improving prior consistency and decoder stability. Results: Extensive experiments demonstrate that LDMAEs significantly outperform state-of-the-art LDM baselines on key metrics—including FID and LPIPS—while reducing sampling computational overhead by over 30%. The framework achieves superior generative quality, inference efficiency, and training stability.
📝 Abstract
In spite of remarkable potential of the Latent Diffusion Models (LDMs) in image generation, the desired properties and optimal design of the autoencoders have been underexplored. In this work, we analyze the role of autoencoders in LDMs and identify three key properties: latent smoothness, perceptual compression quality, and reconstruction quality. We demonstrate that existing autoencoders fail to simultaneously satisfy all three properties, and propose Variational Masked AutoEncoders (VMAEs), taking advantage of the hierarchical features maintained by Masked AutoEncoder. We integrate VMAEs into the LDM framework, introducing Latent Diffusion Models with Masked AutoEncoders (LDMAEs). Through comprehensive experiments, we demonstrate significantly enhanced image generation quality and computational efficiency.