Score
Designs and implements Bayesian methods to propagate uncertainty estimates through connected models, network nodes, or multi-stage pipelines, producing calibrated posterior distributions or end-to-end uncertainty summaries. Builds algorithms that aggregate stage-level uncertainties, compute node-level failure indicators, and translate per-component posterior beliefs into system-level probability estimates for decision-making and risk assessment.
This study addresses the computational bottleneck in evaluating the calibration of nested uncertainty sets within expensive simulation models. To overcome this limitation, the work proposes an efficient calibration assessment method grounded in a Bayesian framework and the Dirichlet-Multinomial model. By exploiting the nested structure, the approach directly processes interval outputs without requiring access to the full predictive distribution. Furthermore, it incorporates Bayes factor testing for statistical inference, substantially reducing the number of independent simulations needed. The proposed method successfully detects model miscalibration in data assimilation tasks under limited simulation budgets. Overall, this work significantly lowers computational costs while demonstrating both the effectiveness and practical utility of the proposed approach for calibrating complex simulation systems.
This work addresses the lack of reliable uncertainty estimation for reasoning failures in agentic retrieval-augmented generation (RAG) systems during multi-hop question answering. It proposes the first uncertainty-aware agentic RAG framework, which integrates semantic disagreement metrics with a generator self-evaluation mechanism to produce stage-wise uncertainty signals. For the first time, a Bayesian network is employed to propagate uncertainty across the system and identify failure-prone components. Experimental results on StrategyQA and HotpotQA demonstrate that the proposed framework significantly improves AUROC and AUARC while reducing ECE and Brier Score, effectively modeling uncertainty accumulation in multi-hop reasoning and enabling both node-level failure warnings and system-level confidence assessment.
In Bayesian inverse problems, surrogate models—constrained by limited simulation budgets and approximation errors—often induce biased parameter estimates and overconfident posterior distributions. To address this, we propose the first scalable Bayesian surrogate modeling framework that rigorously quantifies and propagates uncertainty across the entire pipeline: surrogate construction, posterior inference, and model validation. Methodologically, we integrate Gaussian process surrogate modeling, probabilistic programming, and posterior calibration techniques to design three novel Bayesian inversion algorithms, overcoming classical analytical assumptions and computational bottlenecks. We validate the framework on three real-world linear and nonlinear inverse problems. Results demonstrate substantial improvements in posterior calibration and parameter estimation reliability, leading to reduced decision risk. The framework establishes a new paradigm for robust uncertainty quantification under resource constraints.
This work proposes a unified Bayesian calibration framework to address the lack of reliable and consistent calibration methods for computationally expensive and data-scarce scenarios. The framework uniquely supports both single-output and multi-output complex models within a coherent formulation and is accompanied by ACBICI, a modular open-source Python library. By integrating uncertainty quantification with Bayesian inference, the approach balances usability and extensibility, establishing a closed loop among theory, implementation, and practical application. The study delivers standardized calibration guidelines tailored to real-world engineering challenges and enhances reproducibility and deployment through its open-source toolkit, significantly improving the reliability and accessibility of calibrating complex scientific and engineering models.
This work addresses the challenge of entangled uncertainty sources and the difficulty of disentangling pointwise statistical risk in predictive modeling. We propose a unified generative framework based on approximate Bayesian inference that, for the first time, establishes an explicit, interpretable decomposition linking pointwise statistical risk to two fundamental uncertainty types: aleatoric uncertainty (arising from inherent data noise) and epistemic uncertainty (stemming from model ignorance). The framework jointly generates multiple uncertainty measures while ensuring semantic consistency across them. Experiments on image benchmarks demonstrate significant improvements in out-of-distribution detection and misclassification identification, achieving higher AUROC scores compared to existing methods. Our approach thus provides robust, quantifiable uncertainty estimates essential for downstream uncertainty-aware tasks such as active learning, safe decision-making, and model debugging.
Bayesian inference remains challenging for statisticians and learners due to conceptual ambiguities in its philosophical foundations, difficulties in prior specification, and computational complexity. Method: This paper provides a rigorous yet accessible pedagogical framework for Bayesian inference, systematically integrating core components—including Bayes’ theorem, prior modeling, posterior inference, Bayesian hypothesis testing via Bayes factors, and predictive analysis—while clarifying fundamental distinctions from frequentist paradigms in identifiability, asymptotic theory, and decision-theoretic concepts (e.g., loss functions, credible intervals). It bridges analytical derivations with modern simulation techniques such as MCMC, using canonical statistical models as unifying exemplars. Contribution/Results: The framework innovatively connects foundational concepts to advanced topics—including hierarchical modeling, nonparametric Bayesian methods, and spatiotemporal analysis—and has been successfully applied in political science, network analysis, and spatial statistics, substantially lowering the barrier to learning and applying Bayesian methods in practice.
This work proposes the first general framework to systematically quantify and apportion epistemic uncertainty arising from substituting true subprocesses with approximate or learned submodels in stochastic simulation and digital twin applications. The framework constructs confidence or credible intervals for performance metrics via bootstrapping and Bayesian model averaging, and employs a tree-based decomposition to allocate total output variability to individual submodels, yielding importance scores. It is compatible with both parametric and nonparametric models, supports frequentist and Bayesian paradigms, and accommodates dynamic initialization scenarios. Validation on synthetic data and a call center digital twin demonstrates that the method effectively reveals each submodel’s contribution to overall uncertainty, significantly enhancing the interpretability and reliability of simulation outcomes.
本文提出一种通过概率积分变换和最大均值差异消除重采样噪声的方法,解决了混合不确定性下的模型更新问题,并采用TMCMC进行后验推断。
This work addresses the challenge of surrogate-induced overconfidence in Bayesian inverse problems, where computationally expensive forward models are approximated by surrogates whose uncertainty is often neglected. The authors propose the Expected Posterior (EP) as a principled benchmark for propagating surrogate uncertainty, deriving it for the first time from decision-theoretic and modular Bayesian inference principles. They demonstrate that the commonly used heuristic Expected Utility Posterior (EUP) incurs systematic bias when surrogate uncertainty is non-uniform. To enable practical computation of EP, they develop a randomized kernel-preconditioned Crank–Nicolson (RKpCN) MCMC algorithm, which efficiently approximates the EP even with infinite-dimensional Gaussian process surrogates. This approach significantly enhances the reliability of posterior inference in high-dimensional settings.
本文提出一种基于三向假设检验的框架,用于量化模拟基础推理中的认知校准不确定性,并评估必要的模拟预算。