simulation-based inference

Designs and implements inference workflows that use simulators and neural density estimators to approximate posterior distributions over model parameters when likelihoods are unavailable or intractable; builds neural posterior estimation models, trains them on simulated parameter–data pairs, generates posterior samples and posterior predictive checks, and engineers scalable CPU/GPU training and inference pipelines.

simulation-basedinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Sbi Reloaded: a Toolkit for Simulation-based Inference Workflows

Nov 26, 2024
JB
Jan Boelts
🏛️ University of Tübingen | Tübingen AI Center | TransferLab | appliedAI Institute for Europe | ML Colab | Cluster ML in Science | Google Research | Helmholtz-Zentrum Dresden-Rossendorf | Université Paris-Saclay | INRIA | CEA | Robert Bosch GmbH | School of Informatics | University of Edinburgh | University of Amsterdam | Research and Innovation Center | BMW Group | Institute for Applied Mathematics and Scientific Computing | University of the Bundeswehr Munich | Aix Marseille | INSERM | INS | TU Darmstadt | h

Likelihood-free and gradient-free parameter calibration in black-box simulators poses significant challenges for Bayesian inference. Method: This paper introduces the first simulation-based, fully amortized, gradient-free, and parallelizable neural Bayesian inference framework, accompanied by the open-source PyTorch package SBI. The framework unifies neural posterior estimation (NPE), neural likelihood estimation (NLE), neural ratio estimation (NRE), and mixture density networks (MDNs), integrating Monte Carlo sampling, Bayesian optimization, and simulation scheduling into a modular, end-to-end workflow with production-ready defaults and comprehensive diagnostic tools. Contribution/Results: Evaluated across physics, biology, and astronomy, SBI substantially lowers the barrier to simulation-based inference, accelerates posterior estimation by multiple-fold, and achieves state-of-the-art reusability and scalability.

Enabling Bayesian inference without likelihood evaluationsProviding flexible tools for simulation-based inference workflowsTuning simulator parameters to match observed data

Reducing Calls to the Simulator in Simulation Based Inference (SBI)

Apr 16, 2025
DR
David Refaeli
🏛️ Tel-Aviv University

Likelihood-free inference (LFI) suffers from sample inefficiency when black-box simulators are computationally expensive. To address this, we propose an enhanced Sequential Neural Posterior Estimation (SNPE) framework. Our method introduces two key innovations: (i) integrating a Neural Density Estimator (NDE) as a differentiable likelihood surrogate within the SNPE pipeline, and (ii) adopting Support Points sampling to strategically select informative parameter configurations. This design substantially reduces simulator evaluations while preserving posterior estimation accuracy. Empirical evaluation across multiple benchmark tasks demonstrates that the NDE surrogate exhibits strong stability and generalization. Moreover, Support Points sampling achieves 30–50% fewer simulator calls in several tasks—without degrading, and sometimes even improving, posterior quality. Our approach establishes a new paradigm for efficient Bayesian inference in high-cost simulation settings.

Evaluating Support Points versus random sampling in SNPE algorithmReducing expensive simulator calls in Simulation-Based InferenceUsing Neural Density Estimator surrogate for likelihood sampling

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

Amortized Bayesian Workflow

Sep 06, 2024
MS
Marvin Schmitt

Bayesian inference often faces a trade-off between computational efficiency and posterior accuracy, especially across multiple datasets. This paper proposes an adaptive hybrid inference workflow that—uniquely—integrates amortized variational inference (AVI) with Markov chain Monte Carlo (MCMC) in a dynamically coordinated manner. Leveraging principled posterior diagnostics, it constructs a Pareto frontier to enable automatic, optimal switching between AVI and MCMC. Computational reuse and scheduling optimization further boost inference throughput. The method unifies generative neural network modeling, MCMC refinement, and verifiable diagnostic mechanisms. Evaluated on tens of thousands of real and synthetic datasets, it achieves a 3.2× average speedup over standalone AVI or MCMC baselines, while preserving posterior fidelity—reducing KL divergence by 47% and increasing effective sample size (ESS) by 2.8×. This work delivers a scalable, efficient, and trustworthy solution for large-scale Bayesian inference.

Adaptively choosing inference methods to maintain efficiency and posterior qualityBalancing computational speed and sampling accuracy in Bayesian inferenceIntegrating rapid amortized inference with gold-standard MCMC techniques

Scalable Bayesian Learning with posteriors

May 31, 2024
SD
Samuel Duffield
🏛️ Normal Computing

To address the scalability challenges of Bayesian learning under big data and large models—stemming from high-dimensional posterior approximation—this paper proposes a scalable Bayesian inference framework. Methodologically, it introduces a novel tempered stochastic gradient MCMC perspective, theoretically establishing the asymptotic unbiasedness of deep ensembles. It further provides the first systematic empirical validation of the cold posterior effect in large language models (LLMs), demonstrating improved uncertainty calibration and robustness via Bayesian approximation. Finally, it develops Posteriors, an open-source PyTorch library implementing a unified optimization-and-sampling paradigm, enabling efficient Bayesian inference for models with up to thousands of layers. Experiments across multiple benchmarks and LLM tasks show significant gains in predictive uncertainty calibration and out-of-distribution robustness.

Extensible PyTorch library for Bayesian methodsImproving deep ensembles for unbiased posterior estimationScalable Bayesian learning for high-dimensional models

Latest Papers

What's happening recently
View more

Existing neural posterior estimation methods struggle to handle mixed parameter spaces containing both discrete and continuous variables, limiting their applicability in complex scientific simulations. This work presents the first extension of simulation-based inference (SBI) to such hybrid spaces by introducing a unified joint inference framework: it models discrete parameters via an autoregressive classifier and continuous parameters through a generative model, with both components trained jointly under a single objective. Implemented within the sbi toolkit and accompanied by posterior calibration diagnostics, the proposed method yields accurate and well-calibrated posterior estimates across multiple analytically tractable toy models and realistic scientific simulators, substantially enhancing the practicality and reliability of inference in mixed-parameter settings.

continuous parametersdiscrete parametersmixed parameter spaces

Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.

backpropagation complexityhardware-software co-designheterogeneous accelerators

This work proposes a novel approach to simulation-based inference by integrating large language model–driven program synthesis, enabling joint inference of both model structure and parameters—a capability lacking in traditional methods that rely on fixed, pre-specified simulator architectures. By automatically generating and iteratively refining candidate simulator programs from natural language descriptions, the method transcends rigid modeling assumptions and facilitates the discovery of plausible models directly from open-ended prompts. Empirical evaluations across diverse domains—including deterministic dynamical systems, stochastic epidemic models, and gravitational lensing image analysis—demonstrate its ability to accurately identify data-supported model families, revealing the interplay between the informativeness of observed data and the identifiability of candidate models.

model selectionneural density estimationparameter estimation

This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.

inference pipelinesmicrocontrollerneural inference

This study addresses the high computational cost and insufficient reliability guarantees of neural simulation-based inference by proposing a hybrid inference framework incorporating semiparametric formulations. By introducing two mixture strategies, including latent classes, the method effectively balances inference sensitivity with computational efficiency while preserving the statistical reliability of parametric models. Experimental results demonstrate that the proposed approach significantly reduces computational overhead at the expense of only marginal sensitivity loss. It supports both offline analysis and future trigger-level real-time applications, providing a theoretically grounded and practically valuable solution for efficient and robust statistical inference.

Computational CostLimited-Budget ScenariosNeural Simulation-Based Inference

Hot Scholars

ST

Stefan T. Radev

Assistant Professor, Rensselaer Polytechnic Institute
Deep LearningBayesian StatisticsStochastic ModelsMachine Learning
JR

Jeffrey Regier

Assistant Professor, Department of Statistics, University of Michigan
Bayesian statisticsmachine learningbioinformaticsastronomy
RH

Raphaël Huser

Associate Professor, King Abdullah University of Science and Technology (KAUST)
Statistics of ExtremesSpatio-Temporal StatisticsComputational StatisticsMachine Learning
MS

Marvin Schmitt

ELLIS
Generative Neural NetworksProbabilistic MLUncertainty QuantificationSimulation Intelligence