Score
Train convolutional variational autoencoders: build and optimize encoder–decoder neural models that use convolutional layers with a probabilistic (variational) latent bottleneck to learn compact continuous latent representations and to reconstruct or generate high-dimensional spatial/structured inputs. Work includes selecting convolutional architectures and latent dimensionality, specifying and balancing reconstruction and KL objectives, applying regularization and optimization strategies, and evaluating latent organization such as clustering, disentanglement, and smooth trajectories.
The rapid advancement of generative AI—including GANs, VAEs, and diffusion models—has led to an overwhelming and fragmented literature, necessitating a systematic synthesis. This survey proposes a unified technical taxonomy that integrates the evolutionary trajectories, architectural variants, and hybridization strategies of these three dominant paradigms, clarifying shared optimization principles for generation quality, diversity, and controllability. It introduces, for the first time, a multi-dimensional classification framework spanning model architecture, training mechanisms, and application domains. Furthermore, incorporating ethical considerations and societal impact, the survey identifies three key frontiers: scalability, trustworthy generation, and human-AI collaboration. By unifying conceptual foundations and highlighting emerging challenges, this work delivers a structured, forward-looking technical roadmap for researchers and practitioners in generative AI.
This paper addresses the lack of a unified theoretical framework for variational dimensionality reduction. It proposes a unified Variational Information Bottleneck (VIB) framework that jointly optimizes encoder-based information compression and decoder-based generative fidelity, enabling principled information trade-offs in latent space. Key contributions include: (1) introducing DVSIB and beta-DVCCA—novel methods that extend the multivariate information bottleneck to deep variational settings for the first time; (2) establishing theoretical connections between DSIB and contrastive learning approaches (e.g., Barlow Twins) via mutual information regularization; and (3) proposing symmetric and weighted mutual information regularization to support multi-view representation learning and generative modeling. Evaluated on Noisy MNIST and CIFAR-100, the framework achieves significant improvements in classification accuracy, latent dimension efficiency, and sample efficiency, attaining state-of-the-art or superior performance.
This work addresses the challenge in variational autoencoders (VAEs) of simultaneously achieving high representational capacity and disentangled, low-dimensional latent representations. The authors formulate VAE training as a soft-constrained optimization problem, introducing an entropy-based soft constraint mechanism to regulate the information content of individual latent variables. Coupled with weight filtering, this approach enables automatic pruning of low-entropy dimensions. The proposed method enhances representation efficiency while preserving disentanglement. Experiments demonstrate significant improvements: on dSprites, activation scores increase by 43–62%, FactorVAE score reaches 0.891, and reconstruction error decreases by 38%; on MNIST, over 90% classification accuracy is achieved using only two latent dimensions—reducing input dimensionality by 80% compared to baselines—and training convergence accelerates by 37%.
This paper addresses the fundamental mismatch between the continuous latent space of standard Variational Autoencoders (VAEs) and the inherently discrete nature of data such as text. To resolve this, we propose the Discrete VAE—a VAE explicitly designed for categorical latent variables. Methodologically, we derive the evidence lower bound (ELBO) rigorously from first principles of variational inference under categorical latents and employ the Gumbel-Softmax reparameterization to enable differentiable gradient estimation in discrete latent spaces. Our key contributions are threefold: (1) a tutorial-style, unified theoretical framework for discrete VAEs; (2) a robust and reproducible training paradigm; and (3) publicly released, fully functional code. Experiments demonstrate that the Discrete VAE significantly improves interpretability and structural coherence in discrete data generation, outperforming continuous-latent baselines while preserving principled probabilistic modeling.
This paper addresses the conceptual gap between variational autoencoders (VAEs) and vector-quantized VAEs (VQ-VAEs) in modeling continuous versus discrete latent representations. Methodologically, it proposes a novel autoencoder framework that eliminates both the KL divergence term and the reparameterization trick; instead, it explicitly enforces latent space compactness via learnable clustering centers and employs multi-vector outputs to jointly support continuous interpolation and discrete reconstruction. Key contributions include: (1) uncovering an intrinsic relationship between autoencoder generative fidelity and latent space compactness; (2) establishing a deterministic transition from VAEs to VQ-VAEs without relying on probabilistic assumptions; and (3) empirically validating smooth interpolation and stable reconstruction on MNIST, CelebA, and FashionMNIST. Experiments further reveal that naively increasing the number of output vectors leads to model degradation—manifesting as localized, patchwise discrete encoding—highlighting the critical role of architectural design.
Multimodal variational autoencoders (MVAEs) suffer from overly rigid cross-modal representation coupling, making it difficult to simultaneously ensure high-quality shared representations and modality-specific fidelity. Method: We propose a soft-constrained Mixture-of-Experts (Soft-MoE) prior that replaces hard parameter sharing with learnable gating weights, enabling flexible alignment of modality-specific latent distributions under a unified posterior. This decouples modality-invariant and modality-specific representations while preserving information integrity via variational inference and a soft alignment loss. Contribution/Results: Experiments on multiple benchmarks and real-world multimodal datasets demonstrate that our approach significantly outperforms existing shared-architecture MVAEs. It achieves state-of-the-art performance in both latent representation quality—measured by disentanglement and downstream task accuracy—and missing modality imputation accuracy.
Existing generative approaches typically rely on two-stage training—first pretraining a VAE, then training a generative model in the latent space—resulting in high computational cost and slow sampling. This work proposes CoVAE, a single-stage generative autoencoder framework that integrates consistency modeling into the VAE architecture for end-to-end joint optimization. Its key innovations include: (i) a time-dependent β-scheduling mechanism to control KL divergence, enabling progressive latent-space refinement; (ii) an encoder that explicitly models multi-level noisy latent representations to emulate the diffusion forward process; and (iii) a decoder trained jointly via consistency loss and variational regularization. CoVAE enables high-fidelity generation in a single step or very few steps, significantly outperforming conventional VAEs and existing single-stage methods in both sample quality and inference speed, while offering theoretical coherence and practical deployment efficiency.
Discrete variational autoencoders (VAEs) suffer from non-differentiable discrete latent variables, necessitating biased or high-variance approximations—such as Gumbel-Softmax or REINFORCE—for gradient estimation, which limits high-fidelity image reconstruction. To address this, we propose a reparameterization-free training framework: a nonparametric encoder serves as a proxy to guide the optimization of a parametric encoder via natural gradient updates; integrated with a Transformer-based encoder and an automatic step-size adaptation mechanism, the framework enables end-to-end training. Our key contribution is the first application of natural gradient optimization—originally developed for policy learning—to discrete VAEs, effectively circumventing the bias–variance trade-off inherent in conventional estimators. On ImageNet 256, our method achieves a 20% improvement in Fréchet Inception Distance (FID) over strong baselines, including vector quantized VAEs and Gumbel-Softmax VAEs, demonstrating that high-quality image reconstruction is feasible even under highly compact discrete latent representations.
This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.
This work addresses the challenge of learning disentangled and interpretable latent representations from complex, non-stationary, high-dimensional time-varying signals, which exhibit rich time-frequency structures that conventional variational autoencoders (VAEs) struggle to model effectively. To this end, we propose the Decompositional Variational Autoencoder (DecVAE), which uniquely integrates signal decomposition priors directly into the VAE framework. DecVAE employs an encoder-only architecture and jointly leverages signal decomposition models, contrastive self-supervised tasks, and variational inference to learn multi-subspace latent representations aligned with the intrinsic time-frequency characteristics of the data. Extensive experiments on synthetic data and three scientific datasets demonstrate that DecVAE substantially outperforms existing VAE approaches, achieving significant improvements in disentanglement quality, cross-task generalization, and interpretability of the learned latent representations.
This work addresses the challenge of integrating variational autoencoders (VAEs) as trainable layers within neural networks. It proposes a general framework for flexibly embedding VAEs into arbitrary network architectures, accompanied by an end-to-end training strategy that leverages the reparameterization trick and probabilistic modeling to ensure full differentiability throughout the pipeline. For the first time, this approach enables VAEs to function as plug-and-play modules akin to standard neural network layers, substantially enhancing their compatibility and representational capacity within complex models. Experimental results demonstrate that the proposed VAE layer consistently achieves stable performance across diverse tasks and outperforms conventional standalone VAE models, thereby significantly expanding the applicability of VAEs in deep learning systems.