🤖 AI Summary
This study addresses key limitations in amortized Bayesian inference, where computational constraints induce a representation gap and hinder convergence, often creating an information bottleneck between feature learning and conditional distribution estimation. To overcome these challenges, the authors introduce auxiliary supervised losses to optimize internal representations and propose a general diagnostic mechanism that decouples summary failure from inference failure. Furthermore, an auxiliary guidance strategy is designed to accelerate convergence. Supported by theoretical analysis grounded in the information bottleneck framework, the proposed method significantly improves both convergence speed and estimation accuracy across diverse real-world inference tasks. It effectively mitigates the representation gap and enhances inferential performance under data-scarce conditions.
📝 Abstract
Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal convergence, practical amortized inference still requires iterating over architectures and optimization choices and ultimately ``satisficing'' under finite simulation, compute, and time budgets. Even the best-performing solution may thus retain avoidable representation gaps that typically require problem-specific fixes. Here, we propose a generic alternative which improves training dynamics with auxiliary guidance losses applied to internal representations. Specifically, we show how such guidance leads to faster convergence when training data is abundant and to better performance when it is scarce. We formalize representation gaps as getting stuck in a local optimum at the information bottleneck between the parts of the network tasked with feature learning and those tasked with conditional distribution learning, and offer a generic diagnostic to separate summary failures from inference failures. Finally, we demonstrate that auxiliary supervision improves convergence speed and accuracy on a range of challenging real-world inference problems.