🤖 AI Summary
This work addresses posterior collapse in variational autoencoders (VAEs), a fundamental failure mode in representation learning. Leveraging tools from high-dimensional statistical physics and variational inference theory, we construct a minimal VAE model and rigorously analyze its large-sample asymptotics. We establish, for the first time, the existence of a data-size-independent critical β-value: when the KL-weight β exceeds this threshold, posterior collapse to the prior becomes inevitable. Concurrently, the rate-distortion curve exhibits explicit sample-size dependence, with asymptotic performance improving markedly as data volume grows. Our analysis reveals that the phase transition threshold is jointly governed by β and dataset size, precisely characterizing both collapse conditions and the rate-distortion trade-off. These theoretical predictions are robustly validated across diverse nonlinear VAE architectures and real-world datasets, challenging conventional qualitative interpretations of β-regularization.
📝 Abstract
In variational autoencoders (VAEs), the variational posterior often collapses to the prior, known as posterior collapse, which leads to poor representation learning quality. An adjustable hyperparameter beta has been introduced in VAEs to address this issue. This study sharply evaluates the conditions under which the posterior collapse occurs with respect to beta and dataset size by analyzing a minimal VAE in a high-dimensional limit. Additionally, this setting enables the evaluation of the rate-distortion curve of the VAE. Our results show that, unlike typical regularization parameters, VAEs face"inevitable posterior collapse"beyond a certain beta threshold, regardless of dataset size. Moreover, the dataset-size dependence of the derived rate-distortion curve suggests that relatively large datasets are required to achieve a rate-distortion curve with high rates. These findings robustly explain generalization behavior observed in various real datasets with highly non-linear VAEs.