Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究提出了一种名为metabeta的预训练神经网络,用于解决广义线性混合效应模型(GLMMs)中的贝叶斯推理问题,通过先验摊销和上下文学习方法,实现了比传统NUTS更快更稳定的推理。
📝 Abstract
Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for every new dataset, model and prior. We introduce metabeta, a pretrained neural network for prior-amortized in-context Bayesian inference over GLMMs. Unlike previous neural posterior estimators that fix the prior at training time, metabeta accepts prior families and hyperparameters as inputs at test time, enabling zero-shot generalization. Two set transformers and conditional normalizing flows mirror the posterior's two-level structure (global parameters shared across groups, local parameters per group). The model is trained on millions of realistic simulated datasets spanning continuous, binary, and count outcomes. By default, the flow posterior is refined by Independence Metropolis-Hastings against the unnormalized posterior, so its correctness rests on the sampler rather than the network; this yields tuning-free inference two to three orders of magnitude faster than NUTS. Alternatively, the flow can warm-start NUTS, giving nearly identical inference with substantially increased speed and stability. On controlled benchmarks with ground-truth parameters, metabeta matches NUTS in parameter recovery, calibration and out-of-sample prediction. On out-of-distribution real datasets, its posteriors closely match those of NUTS across all parameter types, and they remain faithful under misspecified likelihoods and priors, out-of-distribution predictors, collinear designs, and data-poor regimes. The model is open-source and open-weights and thus immediately deployable.
Problem

Research questions and friction points this paper is trying to address.

Bayesian Inference
Generalized Linear Mixed-Effects Models
Hierarchical Data
Prior-Amortized
Neural Network
Innovation

Methods, ideas, or system contributions that make the work stand out.

prior-amortized in-context Bayesian inference
generalized linear mixed-effects models
neural posterior estimators
conditional normalizing flows
Independence Metropolis-Hastings