Score
Designs and implements algorithms and evaluation pipelines that compute and validate per-input uncertainty scores for individual model outputs, including methods specialized for diffusion-model denoisers. Typical work includes post-hoc uncertainty estimation on pretrained models, applying Laplace approximations to denoisers, measuring noise-prediction variability along diffusion trajectories, and producing calibrated per-sample uncertainty scores for downstream analysis or monitoring.
Existing 3D molecular diffusion models struggle to provide reliable uncertainty estimates for identifying low-quality generated samples. This work introduces Laplace approximation into the molecular diffusion framework for the first time, proposing a post-hoc, training-free method that quantifies uncertainty by modeling the variability of noise predictions from the denoising network along the generation trajectory. By assigning an uncertainty score to each generated sample and integrating a test-time scaling strategy, the resulting uncertainty estimates exhibit a strong negative correlation with established sample quality metrics. This enables effective filtering of low-quality molecules and substantially enhances the model’s generative performance at inference time without requiring any retraining.
Industrial data-driven models often lack reliable and intrinsically calibrated uncertainty quantification, limiting their deployment in safety-critical applications. This work proposes a diffusion sampling–based posterior inference framework that, for the first time, integrates diffusion probabilistic models into industrial uncertainty quantification. By directly sampling from the posterior predictive distribution, the method yields well-calibrated predictions without requiring post-hoc calibration. It achieves intrinsic calibration while maintaining high predictive accuracy, consistently outperforming existing approaches across synthetic benchmarks, Raman spectroscopy soft sensors, and an industrial ammonia synthesis case study, with simultaneous improvements in both calibration quality and prediction accuracy.
Existing evaluations of Plug-and-Play Diffusion Prior (PnPDP) solvers focus solely on point estimation accuracy, overlooking the stochastic nature of their outputs and the inherent uncertainty of inverse problems, thereby failing to capture the true posterior distribution. This work presents the first uncertainty quantification (UQ)-oriented evaluation framework for PnPDP solvers, introducing a UQ-driven categorization scheme to systematically analyze their behavior at the distributional level. Through toy-model simulations, multiple PnPDP methods, comprehensive UQ metrics, and real-world scientific datasets, experiments validate the efficacy of the proposed classification and reveal distinct uncertainty characteristics across different solvers. The study establishes a novel evaluation paradigm that enables more reliable reconstructions in scientific inverse problems by explicitly accounting for uncertainty.
This work addresses the fundamental trade-off between discrete-time modeling and continuous-time stochastic differential equation (SDE) modeling in denoising diffusion probabilistic models (DDPMs) and score-based generative models (SGMs). Specifically, it tackles the challenge that discretization-induced errors propagate through reverse sampling, degrading sample quality. To this end, we first unify discrete and continuous modeling paradigms by deriving a total variation (TV) distance bound—integrating discrete Girsanov transformation, Pinsker’s inequality, and the data processing inequality. This bound rigorously characterizes the performance limits of both frameworks. We further obtain an analytically tractable TV upper bound that quantitatively exposes the coupled influence of step size, noise schedule, and score estimation error. Our theoretical results provide principled, information-theoretic guidance for designing efficient and robust discrete-time sampling algorithms in diffusion models.
Existing methods for detecting images generated by diffusion models fail to distinguish between aleatoric and epistemic uncertainty, limiting their discriminative performance and generalization capability. This work addresses this limitation by explicitly leveraging epistemic uncertainty for detection—a first in the field—and proposes a Laplace approximation–based approach to estimate epistemic uncertainty in diffusion models. Furthermore, an asymmetric loss function with a large margin is introduced to emphasize the most discriminative components of reconstruction error. The proposed method achieves state-of-the-art performance across multiple large-scale benchmarks and demonstrates significantly improved generalization in detecting images synthesized by previously unseen diffusion models.
This study addresses the miscalibration of Gaussian processes (GPs) in uncertainty quantification (UQ)—specifically, their frequent lack of probabilistic calibration, which undermines convergence in downstream tasks such as Bayesian optimization. To tackle this, we propose Kernel Covariance Validation (KCV), the first framework to systematically exploit the multivariate normal structure of GP predictions for interpretable, computationally tractable calibration diagnostics. KCV integrates multivariate normality testing, uncertainty calibration analysis, and adaptive target design to quantitatively detect and localize model misspecification. Extensive experiments across 1D to high-dimensional GPs demonstrate KCV’s ability to identify canonical calibration failure modes. Results show that KCV significantly improves predictive reliability and enhances both the convergence speed and stability of optimization algorithms. By enabling principled, diagnostic-driven calibration, KCV establishes a new paradigm for trustworthy UQ in GP-based modeling.
This work addresses the lack of theoretical understanding regarding the intrinsic relationship between the noise predictor $varepsilon_ heta(mathbf{x}_t, t)$ and the forward-process noise $varepsilon_t$ in Denoising Diffusion Probabilistic Models (DDPMs). We derive, for the first time, an explicit analytical expression of $varepsilon_ heta(mathbf{x}_t, t)$ in terms of $varepsilon_t$. Leveraging Markov chain modeling, variational inference, and score function theory—combined with stochastic differential equations and the Fokker–Planck equation—we rigorously prove the gradient-log-density identity. This establishes that the learned noise predictor fundamentally implements a weighted conditional expectation of the forward noise along the reverse trajectory. Our analysis provides a rigorous theoretical foundation for the core DDPM objective, significantly enhancing model interpretability and structural transparency.
This work investigates the mechanisms by which diffusion models generate highly realistic images that differ from their training data—referred to as “creativity”—and demonstrates that this capability stems from the alignment between the denoiser architecture and the target data distribution. Through theoretical analysis and empirical experiments, the study derives explicit forms of the generated distribution for linear, polynomial, and bottleneck-style denoisers for the first time, and systematically evaluates the behavior of various architectures, including UNet variants, throughout the diffusion process. The findings reveal that minor architectural modifications to the UNet significantly impact generation fidelity, thereby underscoring the critical role of the denoiser’s inductive bias and its alignment with the target distribution in determining model performance.
Accurately quantifying sample-wise uncertainty in neural networks for high-stakes applications remains challenging, particularly in disentangling epistemic from aleatoric uncertainty—existing additive decomposition methods often fail to do so reliably. Method: We propose a novel uncertainty estimation and decomposition framework grounded in the signal-to-noise ratio (SNR) of class probability distributions. It introduces a variance-gating mechanism to dynamically model prediction reliability and an ensemble confidence factor to adaptively scale probabilistic outputs. Contribution/Results: Our approach reinterprets the “committee diversity collapse” phenomenon and overcomes fundamental limitations of additive decomposition, enabling finer-grained, well-calibrated separation of epistemic and aleatoric uncertainty. Extensive evaluation across multiple benchmark datasets demonstrates superior discriminability and calibration of uncertainty estimates. The framework advances both theoretical understanding and practical deployment of trustworthy AI systems.
Reduced-order models (ROMs) in cloud microphysics lack general, robust uncertainty quantification (UQ) methods. Method: This paper proposes a model-agnostic, plug-and-play conformal prediction framework—the first to systematically apply conformal prediction across the entire latent-space pipeline, including latent dynamics modeling, state reconstruction, and end-to-end forecasting. The method requires no modification to existing ROM architectures or training procedures; it relies solely on offline calibration to produce statistically calibrated prediction intervals for each pipeline component. Contribution/Results: Evaluated on droplet size distribution evolution prediction, the approach significantly improves reliability and coverage accuracy of uncertainty estimates. It achieves well-calibrated UQ across the full pipeline, demonstrating strong generalizability and engineering practicality without sacrificing computational efficiency or model fidelity.