The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degeneration of predictive distributions in wide Bayesian neural networks, where mean-field variational inference becomes dominated by the prior. To mitigate this issue, the work introduces a likelihood tempering mechanism and investigates the evolution of predictive moments in single-hidden-layer linear networks under the infinite-width limit via asymptotic analysis, revealing phase transition behaviors of the mean and variance across different scaling regimes. The primary contributions include deriving exact predictive distributions in the infinite-width limit, establishing critical conditions for temperature decay that counteract degeneracy, and proposing a theoretical framework to recover Neural Network Gaussian Process (NNGP) posterior statistics. Collectively, these results provide rigorous theoretical foundations for reliable inference in high-dimensional Bayesian neural networks.
📝 Abstract
In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to "prior dominance": the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width $M$ grows. Tempering the likelihood, by raising it to the power $1/T$ for a temperature $T < 1$, is equivalent to scaling the KL term by $T$. We ask in this paper how fast $T$ must decrease with $M$ to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form $T = τ/M^{c}$, with constants $τ, c > 0$, as $M \to \infty$ and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at $c = 1/2$, once $τ$ falls below an explicit threshold, and equals the least-squares prediction for $c > 1/2$, whereas the limiting variance keeps its prior value for $c < 1$, matches the NNGP's for $c=1$, and vanishes for $c > 1$. With suitable choices of $τ,c$, one can recover either the NNGP posterior expectation or its variance.
Problem

Research questions and friction points this paper is trying to address.

Bayesian neural networks
variational inference
prior dominance
likelihood tempering
predictive moments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Likelihood Tempering
Variational Bayesian Neural Networks
Phase Transitions
Neural Network Gaussian Process
Prior Dominance
🔎 Similar Papers
No similar papers found.