Latent Diffusion Models with Masked AutoEncoders

📅 2025-07-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing latent diffusion models (LDMs) suffer from a fundamental trade-off among latent-space smoothness, perceptual compression quality, and reconstruction fidelity—largely attributable to suboptimal autoencoder design. Method: We identify this architectural limitation and propose LDMAEs, a novel framework integrating Variational Masked Autoencoders (VMAEs) into the LDM paradigm. LDMAEs are the first to incorporate hierarchical feature modeling—inspired by Masked Autoencoders (MAEs)—into a variational autoencoding structure, enabling joint optimization of smoothness, perceptual quality, and fidelity directly in the latent space. VMAEs employ layered masking and reconstruction to learn compact, perception-driven representations, thereby improving prior consistency and decoder stability. Results: Extensive experiments demonstrate that LDMAEs significantly outperform state-of-the-art LDM baselines on key metrics—including FID and LPIPS—while reducing sampling computational overhead by over 30%. The framework achieves superior generative quality, inference efficiency, and training stability.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersComputer Vision: Diffusion Models for VisionNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
In spite of remarkable potential of the Latent Diffusion Models (LDMs) in image generation, the desired properties and optimal design of the autoencoders have been underexplored. In this work, we analyze the role of autoencoders in LDMs and identify three key properties: latent smoothness, perceptual compression quality, and reconstruction quality. We demonstrate that existing autoencoders fail to simultaneously satisfy all three properties, and propose Variational Masked AutoEncoders (VMAEs), taking advantage of the hierarchical features maintained by Masked AutoEncoder. We integrate VMAEs into the LDM framework, introducing Latent Diffusion Models with Masked AutoEncoders (LDMAEs). Through comprehensive experiments, we demonstrate significantly enhanced image generation quality and computational efficiency.
Problem

Research questions and friction points this paper is trying to address.

Exploring optimal autoencoder design for Latent Diffusion Models
Addressing limitations in latent smoothness and reconstruction quality
Improving image generation efficiency with Variational Masked AutoEncoders
Innovation

Methods, ideas, or system contributions that make the work stand out.

Propose Variational Masked AutoEncoders (VMAEs)
Integrate VMAEs into LDM framework
Enhance image generation quality and efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junho Lee
Seoul National University, Seoul, Korea
J
Jeongwoo Shin
Seoul National University, Seoul, Korea
H
Hyungwook Choi
Seoul National University, Seoul, Korea
Joonseok Lee
Joonseok Lee
Google Research, Seoul National University
Machine LearningComputer VisionVideo UnderstandingRecommendation SystemsCollaborative Filtering