assess extreme-event fidelity and utility

Designs and applies quantitative evaluation methods and diagnostic tests to measure how faithfully models, simulations, or forecast systems reproduce the occurrence, magnitude, timing, and spatial/temporal structure of extreme events. Assesses the practical utility of those outputs for intended decision or analysis purposes by characterizing uncertainties, biases, event detection and attribution performance, and the sensitivity of outcomes to model choices and observational limitations.

assessextreme-eventfidelityand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Assessing Extrapolation of Peaks Over Thresholds with Martingale Testing

Dec 02, 2025
JD
Joseph de Vilmarest
🏛️ Viking Conseil | LPSM | Sorbonne Université

This study addresses the probabilistic extrapolation of sparse extreme precipitation events—those occurring only once or remaining unobserved in the data—by proposing a robust extreme-value-theoretic modeling framework. Methodologically, it employs seasonal decomposition to remove periodic variability; fits the Generalized Pareto Distribution (GPD) to threshold exceedances above high quantiles; introduces a novel martingale-based test to rigorously assess extrapolation reliability; and designs a prior-free, adaptive threshold selection scheme that recasts extreme-value extrapolation as a verifiable stochastic game. Evaluated on the EVA2025 Data Challenge, the framework achieved first place, significantly improving both accuracy and interpretability of extreme-event probability estimates under highly sparse sampling regimes. It establishes a new paradigm for climate risk assessment grounded in statistically principled, transparent, and empirically validated extreme-value inference.

Estimating probability of extreme precipitation eventsExtrapolating extreme values with limited dataSelecting high quantile level using martingale testing

Tail calibration of probabilistic forecasts

Jul 03, 2024
SA
Sam Allen
🏛️ ETH Zurich | University of Bern | KU Leuven | UCLouvain

Existing probabilistic forecasting evaluation methods lack the ability to characterize tail calibration—critical for high-impact extreme events, whose reliability is increasingly vital for risk-informed decision-making. Method: This paper introduces, for the first time, a general definition of tail calibration, rigorously connecting it to classical probabilistic calibration theory and integrating the Peaks-over-Threshold (POT) framework from extreme value theory. We develop an operational diagnostic framework by unifying probabilistic calibration theory, extreme-value statistics, diagnostic statistical tests, and empirical analysis. Contribution/Results: Applied to European precipitation forecasts, our framework significantly improves the quantification of predictive credibility for high-impact, rare events. It enables rigorous assessment of tail behavior in probabilistic forecasts and establishes a novel paradigm for extreme-event risk assessment and decision support.

Assessing tail calibration of probabilistic forecastsConnecting tail calibration to extreme value theoryEvaluating reliability of extreme outcome predictions

This study addresses the lack of interpretable, scalable diagnostic tools for extreme value regression models that can identify regions in covariate space where local fit is poor. The authors propose two visualization-based diagnostics—standardized tail plots and normalized residual plots—leveraging the asymptotic distribution of normalized exceedance probabilities to construct sample-size-invariant uncertainty bounds. This enables consistent assessment of both global and local goodness-of-fit. Notably, the approach provides the first framework for local diagnostics in low-dimensional or non-Euclidean covariate domains, supports model comparison across varying sample sizes, and facilitates large-scale model screening. In two real-world applications, the method successfully evaluated thousands of candidate models, yielding actionable modeling recommendations that substantially enhance the reliability and practical utility of extreme value regression models.

covariate spaceextreme value regressiongoodness-of-fit diagnostics

This study addresses a key limitation in current extreme event attribution practices, which typically estimate attribution metrics separately from observations and models before combining them—a procedure prone to introducing bias. To overcome this, the authors propose a novel parameter-level evidence synthesis framework that directly integrates regression parameters from nonstationary distribution models, thereby transcending the conventional reliance on post-hoc aggregation of attribution metrics. The approach enables unified inference across multiple thresholds and climate scenarios, including counterfactual conditions, by jointly incorporating probability ratio and intensity change metrics. Simulation experiments demonstrate that the method significantly outperforms standard attribution workflows. The framework is successfully applied to attribute the extreme precipitation associated with Storm Boris in September 2024.

climate changeevidence synthesisextreme event attribution

Existing power grid resilience research remains largely conceptual or focuses on isolated components, lacking a system-level, quantifiable definition and assessment framework. Method: Leveraging 15-minute-resolution customer outage time-series data and high-resolution meteorological records, we develop a spatiotemporal statistical model incorporating resilience sensitivity simulation and outage propagation dynamics inference. Contribution/Results: We propose the first system-level, empirically measurable definition of grid resilience. The model uncovers cumulative outage effects under extreme weather, inter-regional outage propagation mechanisms, and systemic response patterns. It identifies critical reinforcement nodes that reduce customer outage magnitude by nearly 50%. Validated across three major U.S. East Coast utility service territories, the model achieves high accuracy in forecasting outage progression—enabling actionable support for real-time dispatch decisions and emergency response.

Analyze large-scale outage data to understand system-level resilienceDevelop predictive model for outage progress during disastersQuantify power grid resilience against extreme weather events

Latest Papers

What's happening recently
View more

Existing predictive evaluation methods struggle to rigorously quantify sampling uncertainty in multidimensional settings, often leading to inflated Type I error rates under multiple comparisons and invalid joint inference. This work proposes a unified statistical framework that constructs simultaneous confidence bands—applicable across multivariate, multi-step-ahead, multi-location, and multi-model configurations—for joint inference on mean, quantile, and distributional forecasts. The approach builds upon a multivariate extension of the Diebold–Mariano test and incorporates bootstrap-based uncertainty quantification. Empirical validation in macroeconomic and weather forecasting applications demonstrates the framework’s ability to effectively discern predictive advantages of time-varying parameter models against data-driven alternatives.

Forecast ComparisonJoint InferenceMultiple Comparisons

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This study addresses the limitation of traditional probabilistic forecasting in expressing epistemic uncertainty—specifically, the inability to convey “I don’t know”—due to normalization constraints. By leveraging possibility theory, the authors propose a novel framework that explicitly models ignorance as the non-normalization of possibility distributions, integrating perspectives from possibility, probability, and classification. The framework enables fine-grained diagnosis of forecast failure modes through a five-dimensional scoring system—assessing validity, sharpness, ignorance, dominance, and an overall composite metric—augmented with operational metrics such as POD, FAR, and CSI. Experiments on Storm Prediction Center convective outlook data demonstrate that explicitly representing ignorance yields superior performance compared to enforcing normalization, and synthetic reforecast experiments validate both the method’s efficacy and its interpretability through visualization.

forecast verificationignorancepossibilistic forecasts

This study addresses the challenge of modeling rare, extreme anomalies in aircraft manufacturing, which often exhibit heavy-tailed distributions and spatial dependencies among multiple outputs—features poorly captured by conventional machine learning methods. To this end, the authors propose a novel extreme-value spatial model that integrates extreme value theory with multi-output spatial modeling. The approach employs a bilinear function to characterize dynamic interactions between control variables and measurement locations across two spatial domains and introduces a graph-assisted composite likelihood estimation method to handle high-dimensional outputs. The resulting framework jointly models marginal extreme behaviors and extremal dependence structures among outputs, supported by an efficient computational algorithm. Experiments on a composite-material aircraft production system demonstrate that the proposed method significantly outperforms existing techniques in predicting extreme events, thereby enhancing quality control and operational safety.

extremal dependenceextreme eventsheavy-tailed distributions

This study addresses the challenge of fairly evaluating artificial intelligence (AI) models against physics-based numerical weather prediction (NWP) systems in forecasting extreme weather events. To this end, the authors propose a weighted version of the continuous ranked probability score in latent space (Weighted PCRPS), coupled with Isotonic Distributional Regression (IDR) to convert deterministic forecasts into probabilistic predictions for rigorous assessment. The weighting scheme emphasizes extreme events, while IDR ensures evaluation fairness through its optimality properties. Comprehensive comparisons on the WeatherBench 2 dataset among leading AI models—including GraphCast, Pangu-Weather, and FuXi—and ECMWF’s high-resolution NWP system demonstrate that FuXi achieves overall superior performance across key extreme meteorological variables such as pressure, temperature, wind speed, and precipitation, highlighting the potential of AI-based approaches to surpass traditional NWP in extreme weather forecasting.

AI weather predictionextreme weather eventsfair comparison