$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
This study addresses the prohibitive computational cost of selecting prior precision in Laplace approximations for large neural networks. By deriving a prior covariance rescaling mechanism grounded in maximal update parameterization ($\mu$P), this work proposes $\sigma$Transfer, a hyperparameter-free method that zero-shot transfers prior precision from small to large models, with theoretical guarantees that posterior stability converges as network width increases. Empirically, $\sigma$Transfer achieves an approximately 5000-fold speedup on MNIST with only a 0.002 increase in negative log-likelihood (NLL). Furthermore, it enables efficient precision transfer across models scaling from 1B to 7B parameters, yielding NLL degradation below $10^{-4}$. These results establish a scalable, low-cost paradigm for uncertainty estimation in large-scale models.