Multivariate Variational Autoencoder

📅 2025-11-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional VAEs restrict the posterior to diagonal covariance structures, limiting their ability to capture correlations among latent variables. To address this, we propose the Full-Covariance Variational Autoencoder (FC-VAE), enabling efficient full-covariance inference while preserving analytical tractability of Gaussian posteriors. Our key innovation is a structured decomposition of the posterior covariance: ( L = C cdot ext{diag}(sigma) ), where ( C ) is a global low-rank coupling matrix and ( ext{diag}(sigma) ) is a sample-specific diagonal scaling matrix. This formulation permits closed-form KL divergence computation and standard reparameterization without auxiliary approximations or sampling. FC-VAE significantly improves reconstruction accuracy (MSE), probabilistic calibration (NLL, Brier score, ECE), and unsupervised clustering performance (NMI, ARI), especially at moderate latent dimensions. Extensive experiments across multiple image datasets validate its effectiveness and generalization capability.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersComputer Vision: Learning & Optimization for CVReasoning under Uncertainty: Relational Probabilistic Models

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
We present the Multivariate Variational Autoencoder (MVAE), a VAE variant that preserves Gaussian tractability while lifting the diagonal posterior restriction. MVAE factorizes each posterior covariance, where a emph{global} coupling matrix $mathbf{C}$ induces dataset-wide latent correlations and emph{per-sample} diagonal scales modulate local uncertainty. This yields a full-covariance family with analytic KL and an efficient reparameterization via $mathbf{L}=mathbf{C}mathrm{diag}(oldsymbol{sigma})$. Across Larochelle-style MNIST variants, Fashion-MNIST, CIFAR-10, and CIFAR-100, MVAE consistently matches or improves reconstruction (MSE~$downarrow$) and delivers robust gains in calibration (NLL/Brier/ECE~$downarrow$) and unsupervised structure (NMI/ARI~$uparrow$) relative to diagonal-covariance VAEs with matched capacity, especially at mid-range latent sizes. Latent-plane visualizations further indicate smoother, more coherent factor traversals and sharper local detail. We release a fully reproducible implementation with training/evaluation scripts and sweep utilities to facilitate fair comparison and reuse.
Problem

Research questions and friction points this paper is trying to address.

Lifting diagonal posterior restriction in VAEs
Modeling dataset-wide latent correlations and local uncertainty
Improving reconstruction, calibration, and unsupervised structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

MVAE uses global coupling matrix for latent correlations
It employs per-sample diagonal scales for local uncertainty
This enables full-covariance Gaussian posteriors with analytic KL
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Işık University
M
Mehmet Can Yavuz
Faculty of Engineering and Natural Sciences, Işık University, İstanbul, Türkiye