🤖 AI Summary
Traditional VAEs restrict the posterior to diagonal covariance structures, limiting their ability to capture correlations among latent variables. To address this, we propose the Full-Covariance Variational Autoencoder (FC-VAE), enabling efficient full-covariance inference while preserving analytical tractability of Gaussian posteriors. Our key innovation is a structured decomposition of the posterior covariance: ( L = C cdot ext{diag}(sigma) ), where ( C ) is a global low-rank coupling matrix and ( ext{diag}(sigma) ) is a sample-specific diagonal scaling matrix. This formulation permits closed-form KL divergence computation and standard reparameterization without auxiliary approximations or sampling. FC-VAE significantly improves reconstruction accuracy (MSE), probabilistic calibration (NLL, Brier score, ECE), and unsupervised clustering performance (NMI, ARI), especially at moderate latent dimensions. Extensive experiments across multiple image datasets validate its effectiveness and generalization capability.
📝 Abstract
We present the Multivariate Variational Autoencoder (MVAE), a VAE variant that preserves Gaussian tractability while lifting the diagonal posterior restriction. MVAE factorizes each posterior covariance, where a emph{global} coupling matrix $mathbf{C}$ induces dataset-wide latent correlations and emph{per-sample} diagonal scales modulate local uncertainty. This yields a full-covariance family with analytic KL and an efficient reparameterization via $mathbf{L}=mathbf{C}mathrm{diag}(oldsymbol{sigma})$. Across Larochelle-style MNIST variants, Fashion-MNIST, CIFAR-10, and CIFAR-100, MVAE consistently matches or improves reconstruction (MSE~$downarrow$) and delivers robust gains in calibration (NLL/Brier/ECE~$downarrow$) and unsupervised structure (NMI/ARI~$uparrow$) relative to diagonal-covariance VAEs with matched capacity, especially at mid-range latent sizes. Latent-plane visualizations further indicate smoother, more coherent factor traversals and sharper local detail. We release a fully reproducible implementation with training/evaluation scripts and sweep utilities to facilitate fair comparison and reuse.