🤖 AI Summary
This study addresses the limitation in sampling quality caused by the mismatch between prior and posterior distributions in variational autoencoders (VAEs) for continuous sequence generation. To overcome this, we propose a general generative framework based on an empirical autoregressive latent prior, replacing the standard Gaussian assumption with a data-driven prior and enabling efficient variational inference through a single linear layer. This architecture significantly mitigates discrepancies in the latent space distribution, thereby enhancing generative fidelity. Experimental results demonstrate that the proposed approach achieves synthesis quality comparable to diffusion models on image and audio generation tasks while offering substantially faster inference speeds.
📝 Abstract
We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the self-predicted priors, EVA significantly alleviates the latent distribution gap between prior and posterior which is typically observed in conventional VAEs, and leads to high-fidelity ancestral sampling for sequential data generation. Extensive experiments on image and sound synthesis demonstrate that EVA achieves competitive generation quality with autoregressive diffusion baselines despite its much faster inference time.