π€ AI Summary
This study addresses the unclear optimal allocation of parameter stochasticity in Bayesian neural networks, highlighting the need to balance approximation capacity with computational efficiency. We propose a prior-scale-based deep weight decomposition method that achieves stochasticity sparsification by fitting Gaussian process priors via maximum mean discrepancy and employing a threshold-switching mechanism to convert low-scale parameters into deterministic optimization. Furthermore, we introduce a general density approximation certificate verifiable in linear time, revealing that hybrid schemes fundamentally constitute Type-II MAP stochastic approximation. Experiments demonstrate that the proposed approach attains fully stochastic network performance on UCI benchmarks using only half the deterministic parameters, while significantly outperforming existing baselines on bimodal tasks.
π Abstract
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.