Prior laundering: learned priors with inherited, undetectable overconfidence

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical yet previously unexamined issue in data-scarce inverse problems—such as those in seismic and medical imaging—where generative priors trained on historical reconstruction archives inadvertently encode prior knowledge rather than observational fidelity. The resulting posterior uncertainty in Bayesian inversion reflects this archival prior rather than the actual confidence warranted by new measurements, a phenomenon we term “prior laundering.” We formally characterize this mechanism, showing it is equivalent to performing a single EM update on an outdated regularizer, which leads to insufficient coverage of blind subspaces and overconfident reconstructions. Through a synthesis of generative modeling (diffusion models and normalizing flows), Bayesian inverse problem theory, and simulation-based calibration, experiments demonstrate that archive-trained models exhibit significantly reduced coverage in operator nullspaces, whereas models trained on genuine data behave as expected—highlighting both the severity and undetectability of this pitfall in deployment.
📝 Abstract
Learned generative priors are increasingly used for ill-posed Bayesian inverse problems, their posterior uncertainty treated as earned from data. But training one requires truths, scarce in seismic and medical imaging, so the recourse is an archive of legacy reconstructions---prior laundering. Where the measurements are uninformative the posterior reverts to the prior, so the uncertainty reported there is the archive's, not the data's, and nothing in deployment reveals it: truths differing only on those directions induce identical data laws, and self-consistency checks such as simulation-based calibration pass whatever the prior believes. That belief has an exact source: averaging the legacy posterior over the measurements yields the old regularizer advanced a single expectation--maximization step---improved where the data resolve, frozen where they cannot. It is overconfident wherever the inherited belief is tighter than the truth. A single-best archive is worse, collapsing the blind credible interval to zero width. Deployed, a diffusion prior fit to the archive under-covers the operator's blind subspace, unlike a truth-trained control, and a normalizing flow does the same on a nonlinear groundwater operator. We recommend reporting which directions the measurements resolve, separating confidence the data support from belief inherited through the pipeline.
Problem

Research questions and friction points this paper is trying to address.

prior laundering
Bayesian inverse problems
overconfidence
posterior uncertainty
ill-posed problems
Innovation

Methods, ideas, or system contributions that make the work stand out.

prior laundering
Bayesian inverse problems
overconfidence
generative priors
uncertainty quantification