Score
Designs and applies quantitative evaluation methods and diagnostic tests to measure how faithfully models, simulations, or forecast systems reproduce the occurrence, magnitude, timing, and spatial/temporal structure of extreme events. Assesses the practical utility of those outputs for intended decision or analysis purposes by characterizing uncertainties, biases, event detection and attribution performance, and the sensitivity of outcomes to model choices and observational limitations.
Extreme events—such as stock market crashes, earthquakes, and pandemics—are rare, catastrophic, and exhibit system-wide propagation, leading to severe data scarcity that undermines data-driven modeling. To address this, we present the first systematic survey of synthetic data generation methods tailored to extreme events and propose the first dedicated generative framework for extremely rare events. We design a customized evaluation suite encompassing statistical fidelity, dependency preservation, visual plausibility, and task-oriented utility, rigorously analyzing metric validity under heavy-tailed distributions. Our framework unifies generative models (GANs, diffusion models, VAEs), large language models, statistical modeling, and targeted resampling strategies. We curate benchmark datasets across finance, meteorology, geoscience, and epidemiology, identifying underexplored domains—including behavioral finance, wildfire dynamics, and windstorm modeling—and distill key open challenges to advance the reliability and practicality of extreme-event modeling.
This study addresses the probabilistic extrapolation of sparse extreme precipitation events—those occurring only once or remaining unobserved in the data—by proposing a robust extreme-value-theoretic modeling framework. Methodologically, it employs seasonal decomposition to remove periodic variability; fits the Generalized Pareto Distribution (GPD) to threshold exceedances above high quantiles; introduces a novel martingale-based test to rigorously assess extrapolation reliability; and designs a prior-free, adaptive threshold selection scheme that recasts extreme-value extrapolation as a verifiable stochastic game. Evaluated on the EVA2025 Data Challenge, the framework achieved first place, significantly improving both accuracy and interpretability of extreme-event probability estimates under highly sparse sampling regimes. It establishes a new paradigm for climate risk assessment grounded in statistically principled, transparent, and empirically validated extreme-value inference.
Existing probabilistic forecasting evaluation methods lack the ability to characterize tail calibration—critical for high-impact extreme events, whose reliability is increasingly vital for risk-informed decision-making. Method: This paper introduces, for the first time, a general definition of tail calibration, rigorously connecting it to classical probabilistic calibration theory and integrating the Peaks-over-Threshold (POT) framework from extreme value theory. We develop an operational diagnostic framework by unifying probabilistic calibration theory, extreme-value statistics, diagnostic statistical tests, and empirical analysis. Contribution/Results: Applied to European precipitation forecasts, our framework significantly improves the quantification of predictive credibility for high-impact, rare events. It enables rigorous assessment of tail behavior in probabilistic forecasts and establishes a novel paradigm for extreme-event risk assessment and decision support.
This study addresses the lack of interpretable, scalable diagnostic tools for extreme value regression models that can identify regions in covariate space where local fit is poor. The authors propose two visualization-based diagnostics—standardized tail plots and normalized residual plots—leveraging the asymptotic distribution of normalized exceedance probabilities to construct sample-size-invariant uncertainty bounds. This enables consistent assessment of both global and local goodness-of-fit. Notably, the approach provides the first framework for local diagnostics in low-dimensional or non-Euclidean covariate domains, supports model comparison across varying sample sizes, and facilitates large-scale model screening. In two real-world applications, the method successfully evaluated thousands of candidate models, yielding actionable modeling recommendations that substantially enhance the reliability and practical utility of extreme value regression models.
This study addresses a key limitation in current extreme event attribution practices, which typically estimate attribution metrics separately from observations and models before combining them—a procedure prone to introducing bias. To overcome this, the authors propose a novel parameter-level evidence synthesis framework that directly integrates regression parameters from nonstationary distribution models, thereby transcending the conventional reliance on post-hoc aggregation of attribution metrics. The approach enables unified inference across multiple thresholds and climate scenarios, including counterfactual conditions, by jointly incorporating probability ratio and intensity change metrics. Simulation experiments demonstrate that the method significantly outperforms standard attribution workflows. The framework is successfully applied to attribute the extreme precipitation associated with Storm Boris in September 2024.
Existing power grid resilience research remains largely conceptual or focuses on isolated components, lacking a system-level, quantifiable definition and assessment framework. Method: Leveraging 15-minute-resolution customer outage time-series data and high-resolution meteorological records, we develop a spatiotemporal statistical model incorporating resilience sensitivity simulation and outage propagation dynamics inference. Contribution/Results: We propose the first system-level, empirically measurable definition of grid resilience. The model uncovers cumulative outage effects under extreme weather, inter-regional outage propagation mechanisms, and systemic response patterns. It identifies critical reinforcement nodes that reduce customer outage magnitude by nearly 50%. Validated across three major U.S. East Coast utility service territories, the model achieves high accuracy in forecasting outage progression—enabling actionable support for real-time dispatch decisions and emergency response.
Existing predictive evaluation methods struggle to rigorously quantify sampling uncertainty in multidimensional settings, often leading to inflated Type I error rates under multiple comparisons and invalid joint inference. This work proposes a unified statistical framework that constructs simultaneous confidence bands—applicable across multivariate, multi-step-ahead, multi-location, and multi-model configurations—for joint inference on mean, quantile, and distributional forecasts. The approach builds upon a multivariate extension of the Diebold–Mariano test and incorporates bootstrap-based uncertainty quantification. Empirical validation in macroeconomic and weather forecasting applications demonstrates the framework’s ability to effectively discern predictive advantages of time-varying parameter models against data-driven alternatives.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the limitation of traditional probabilistic forecasting in expressing epistemic uncertainty—specifically, the inability to convey “I don’t know”—due to normalization constraints. By leveraging possibility theory, the authors propose a novel framework that explicitly models ignorance as the non-normalization of possibility distributions, integrating perspectives from possibility, probability, and classification. The framework enables fine-grained diagnosis of forecast failure modes through a five-dimensional scoring system—assessing validity, sharpness, ignorance, dominance, and an overall composite metric—augmented with operational metrics such as POD, FAR, and CSI. Experiments on Storm Prediction Center convective outlook data demonstrate that explicitly representing ignorance yields superior performance compared to enforcing normalization, and synthetic reforecast experiments validate both the method’s efficacy and its interpretability through visualization.
This study addresses the challenge of modeling rare, extreme anomalies in aircraft manufacturing, which often exhibit heavy-tailed distributions and spatial dependencies among multiple outputs—features poorly captured by conventional machine learning methods. To this end, the authors propose a novel extreme-value spatial model that integrates extreme value theory with multi-output spatial modeling. The approach employs a bilinear function to characterize dynamic interactions between control variables and measurement locations across two spatial domains and introduces a graph-assisted composite likelihood estimation method to handle high-dimensional outputs. The resulting framework jointly models marginal extreme behaviors and extremal dependence structures among outputs, supported by an efficient computational algorithm. Experiments on a composite-material aircraft production system demonstrate that the proposed method significantly outperforms existing techniques in predicting extreme events, thereby enhancing quality control and operational safety.
This study addresses the challenge of fairly evaluating artificial intelligence (AI) models against physics-based numerical weather prediction (NWP) systems in forecasting extreme weather events. To this end, the authors propose a weighted version of the continuous ranked probability score in latent space (Weighted PCRPS), coupled with Isotonic Distributional Regression (IDR) to convert deterministic forecasts into probabilistic predictions for rigorous assessment. The weighting scheme emphasizes extreme events, while IDR ensures evaluation fairness through its optimality properties. Comprehensive comparisons on the WeatherBench 2 dataset among leading AI models—including GraphCast, Pangu-Weather, and FuXi—and ECMWF’s high-resolution NWP system demonstrate that FuXi achieves overall superior performance across key extreme meteorological variables such as pressure, temperature, wind speed, and precipitation, highlighting the potential of AI-based approaches to surpass traditional NWP in extreme weather forecasting.