🤖 AI Summary
This study addresses the prohibitive computational cost of selecting prior precision in Laplace approximations for large neural networks. By deriving a prior covariance rescaling mechanism grounded in maximal update parameterization ($\mu$P), this work proposes $\sigma$Transfer, a hyperparameter-free method that zero-shot transfers prior precision from small to large models, with theoretical guarantees that posterior stability converges as network width increases. Empirically, $\sigma$Transfer achieves an approximately 5000-fold speedup on MNIST with only a 0.002 increase in negative log-likelihood (NLL). Furthermore, it enables efficient precision transfer across models scaling from 1B to 7B parameters, yielding NLL degradation below $10^{-4}$. These results establish a scalable, low-cost paradigm for uncertainty estimation in large-scale models.
📝 Abstract
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization ($μ\mathrm{P}$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to $σ\mathrm{Transfer}$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σ\mathrm{Transfer}$ across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.