measure ensemble uncertainty

Design and implement quantitative metrics and analyses to characterize uncertainty in ensemble model outputs, including interval widths, ensemble variance and spread, and measures of calibration and sharpness. Perform spread–skill analyses and report correlations among uncertainty metrics to evaluate how ensemble spread relates to predictive error and overall reliability.

measureensembleuncertainty

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.

Machine LearningModel EvaluationStability and Reliability

Current performance evaluation metrics—such as accuracy and F1 score—are typically reported as point estimates, ignoring the uncertainty induced by data clustering structures. This oversight often leads to underestimation of variability and potentially misleading model comparisons. To address this, this work proposes a unified framework that expresses a broad class of performance metrics as smooth functionals of the confusion matrix probabilities. By integrating a cluster-robust sandwich variance estimator, the framework enables valid confidence interval construction, hypothesis testing, and paired model comparison. It represents the first systematic application of cluster-robust inference to predictive performance evaluation, accommodating both binary and multiclass settings, and further provides asymptotic theory–based methods for power and sample size calculations. Simulations demonstrate that the proposed approach achieves near-nominal coverage across diverse dependence structures and substantially outperforms conventional methods that ignore clustering; real-data analyses confirm that accounting for clustering can materially alter evaluation conclusions.

clustered datadependent datamodel evaluation

Ensemble learning for predictive uncertainty estimation with application to the correction of satellite precipitation products

Mar 14, 2024
GP
Georgia Papacharalampous
🏛️ National Technical University of Athens

This study addresses the challenge of calibrating satellite-based precipitation products to improve uncertainty quantification in probabilistic spatial rainfall forecasting. We propose an ensemble framework integrating multiple quantile regression models, systematically evaluating nine quantile learners—including Quantile Regression (QR), Quantile Random Forest (QRF), Generalized Random Forest (GRF), Gradient Boosting Machine (GBM), LightGBM, and Quantile Recurrent Neural Network (QRNN)—alongside six ensemble strategies and three simple combination methods. A novel lightweight feature engineering approach is introduced, incorporating geographic distance-weighted satellite precipitation and elevation features. Experiments on 15 years of monthly data across the Continental United States (CONUS) demonstrate that QR- and QRNN-based ensembles achieve average improvements of 3.91%–8.95% in predictive accuracy within the 0.025–0.975 quantile interval, significantly outperforming single-model quantile regression baselines. These results validate the efficacy and robustness of multi-model quantile ensembling for probabilistic precipitation prediction.

Quantile RegressionRainfall PredictionSatellite Data

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.

Measuring individual model importance in ensemble forecasting accuracyProposing practical methods to assess model contribution to ensemble performanceRevealing unique model features beyond standard accuracy metrics

Latest Papers

What's happening recently
View more

This study addresses the common issue in ensemble forecasting wherein insufficiently rapid growth of ensemble spread leads to inadequate representation of uncertainty. Using the Lorenz '96 system, the work systematically disentangles intrinsic variability, initial condition perturbations, and stochastic model uncertainty to evaluate how various ensemble configurations and parameterization schemes influence spread evolution. It introduces novel Bayesian and streaming stochastic parameterizations featuring temporally coherent structures, revealing that perturbations primarily govern the rate of trajectory decorrelation rather than long-term variance. The analysis further elucidates the interaction mechanisms among distinct uncertainty sources. Experimental results demonstrate that the proposed methods significantly enhance early spread growth and improve consistency between ensemble spread and forecast error, thereby offering theoretical insights and practical guidance for uncertainty modeling in numerical weather prediction systems.

chaotic dynamicsensemble spreadforecast error

Traditional confidence interval plots in multi-model climate prediction visualization often obscure individual model characteristics, leading users to misinterpret the underlying distribution—frequently assuming normality where none exists. To address this, this work proposes a Weighted Multi-Forecast Visualization (MFV) approach that leverages visual variables such as line width and opacity, combined with a downsampling strategy, to preserve accurate perception of the true predictive distribution while effectively conveying additional forecast attributes. Through a preregistered experimental design and large-scale user study, results demonstrate that MFV significantly improves users’ accuracy in identifying predictive distributions and reduces erroneous assumptions of normality. Moreover, the weighted MFV variant successfully overcomes the expressive limitations of conventional summary-based visualizations without compromising distributional fidelity.

climate forecastforecast clutterforecast distribution

Existing predictive evaluation methods struggle to rigorously quantify sampling uncertainty in multidimensional settings, often leading to inflated Type I error rates under multiple comparisons and invalid joint inference. This work proposes a unified statistical framework that constructs simultaneous confidence bands—applicable across multivariate, multi-step-ahead, multi-location, and multi-model configurations—for joint inference on mean, quantile, and distributional forecasts. The approach builds upon a multivariate extension of the Diebold–Mariano test and incorporates bootstrap-based uncertainty quantification. Empirical validation in macroeconomic and weather forecasting applications demonstrates the framework’s ability to effectively discern predictive advantages of time-varying parameter models against data-driven alternatives.

Forecast ComparisonJoint InferenceMultiple Comparisons

This study addresses the insufficient quantification of uncertainty in numerical weather prediction by introducing, for the first time, multiple conformal prediction (CP) methods—including standard CP, normalized CP, and conformal quantile regression—into a one-dimensional shallow-water model data assimilation framework. These methods are integrated with the ensemble Kalman filter to produce prediction intervals endowed with finite-sample theoretical guarantees. Systematic evaluation using metrics such as average coverage, interval width, upper and lower tail miss rates, and interval score demonstrates that CP effectively characterizes forecast uncertainty and complements traditional ensemble-based approaches within the data assimilation cycle. The work establishes a novel paradigm and empirical foundation for uncertainty quantification that synergistically combines machine learning with physics-driven modeling.

conformal predictiondata assimilationensemble methods

This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.

CalibrationCross-Project PredictionPerformance Evaluation

Hot Scholars

JS

John S. Schreck

National Center for Atmospheric Research (NCAR)
nanotechnologystatistical mechanicsmolecular simulationmachine learning
CL

Chenning Li

PhD student at MIT CSAIL
Network SimulationsML Systems
LW

Lu Wang

Associate Professor, Computer Science and Engineering, University of Michigan
Natural Language ProcessingComputational Social ScienceMachine Learning
YZ

Yunxiang Zhang

University of Michigan
Natural Language Processing