Score
Designs and implements inference workflows that use simulators and neural density estimators to approximate posterior distributions over model parameters when likelihoods are unavailable or intractable; builds neural posterior estimation models, trains them on simulated parameter–data pairs, generates posterior samples and posterior predictive checks, and engineers scalable CPU/GPU training and inference pipelines.
Likelihood-free and gradient-free parameter calibration in black-box simulators poses significant challenges for Bayesian inference. Method: This paper introduces the first simulation-based, fully amortized, gradient-free, and parallelizable neural Bayesian inference framework, accompanied by the open-source PyTorch package SBI. The framework unifies neural posterior estimation (NPE), neural likelihood estimation (NLE), neural ratio estimation (NRE), and mixture density networks (MDNs), integrating Monte Carlo sampling, Bayesian optimization, and simulation scheduling into a modular, end-to-end workflow with production-ready defaults and comprehensive diagnostic tools. Contribution/Results: Evaluated across physics, biology, and astronomy, SBI substantially lowers the barrier to simulation-based inference, accelerates posterior estimation by multiple-fold, and achieves state-of-the-art reusability and scalability.
Likelihood-free inference (LFI) suffers from sample inefficiency when black-box simulators are computationally expensive. To address this, we propose an enhanced Sequential Neural Posterior Estimation (SNPE) framework. Our method introduces two key innovations: (i) integrating a Neural Density Estimator (NDE) as a differentiable likelihood surrogate within the SNPE pipeline, and (ii) adopting Support Points sampling to strategically select informative parameter configurations. This design substantially reduces simulator evaluations while preserving posterior estimation accuracy. Empirical evaluation across multiple benchmark tasks demonstrates that the NDE surrogate exhibits strong stability and generalization. Moreover, Support Points sampling achieves 30–50% fewer simulator calls in several tasks—without degrading, and sometimes even improving, posterior quality. Our approach establishes a new paradigm for efficient Bayesian inference in high-cost simulation settings.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
Bayesian inference often faces a trade-off between computational efficiency and posterior accuracy, especially across multiple datasets. This paper proposes an adaptive hybrid inference workflow that—uniquely—integrates amortized variational inference (AVI) with Markov chain Monte Carlo (MCMC) in a dynamically coordinated manner. Leveraging principled posterior diagnostics, it constructs a Pareto frontier to enable automatic, optimal switching between AVI and MCMC. Computational reuse and scheduling optimization further boost inference throughput. The method unifies generative neural network modeling, MCMC refinement, and verifiable diagnostic mechanisms. Evaluated on tens of thousands of real and synthetic datasets, it achieves a 3.2× average speedup over standalone AVI or MCMC baselines, while preserving posterior fidelity—reducing KL divergence by 47% and increasing effective sample size (ESS) by 2.8×. This work delivers a scalable, efficient, and trustworthy solution for large-scale Bayesian inference.
To address the scalability challenges of Bayesian learning under big data and large models—stemming from high-dimensional posterior approximation—this paper proposes a scalable Bayesian inference framework. Methodologically, it introduces a novel tempered stochastic gradient MCMC perspective, theoretically establishing the asymptotic unbiasedness of deep ensembles. It further provides the first systematic empirical validation of the cold posterior effect in large language models (LLMs), demonstrating improved uncertainty calibration and robustness via Bayesian approximation. Finally, it develops Posteriors, an open-source PyTorch library implementing a unified optimization-and-sampling paradigm, enabling efficient Bayesian inference for models with up to thousands of layers. Experiments across multiple benchmarks and LLM tasks show significant gains in predictive uncertainty calibration and out-of-distribution robustness.
Existing neural posterior estimation methods struggle to handle mixed parameter spaces containing both discrete and continuous variables, limiting their applicability in complex scientific simulations. This work presents the first extension of simulation-based inference (SBI) to such hybrid spaces by introducing a unified joint inference framework: it models discrete parameters via an autoregressive classifier and continuous parameters through a generative model, with both components trained jointly under a single objective. Implemented within the sbi toolkit and accompanied by posterior calibration diagnostics, the proposed method yields accurate and well-calibrated posterior estimates across multiple analytically tractable toy models and realistic scientific simulators, substantially enhancing the practicality and reliability of inference in mixed-parameter settings.
Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.
This work proposes a novel approach to simulation-based inference by integrating large language model–driven program synthesis, enabling joint inference of both model structure and parameters—a capability lacking in traditional methods that rely on fixed, pre-specified simulator architectures. By automatically generating and iteratively refining candidate simulator programs from natural language descriptions, the method transcends rigid modeling assumptions and facilitates the discovery of plausible models directly from open-ended prompts. Empirical evaluations across diverse domains—including deterministic dynamical systems, stochastic epidemic models, and gravitational lensing image analysis—demonstrate its ability to accurately identify data-supported model families, revealing the interplay between the informativeness of observed data and the identifiability of candidate models.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
This study addresses the high computational cost and insufficient reliability guarantees of neural simulation-based inference by proposing a hybrid inference framework incorporating semiparametric formulations. By introducing two mixture strategies, including latent classes, the method effectively balances inference sensitivity with computational efficiency while preserving the statistical reliability of parametric models. Experimental results demonstrate that the proposed approach significantly reduces computational overhead at the expense of only marginal sensitivity loss. It supports both offline analysis and future trigger-level real-time applications, providing a theoretically grounded and practically valuable solution for efficient and robust statistical inference.