Score
Designs and trains probabilistic ensemble prediction systems whose training objective and loss functions are based on the Continuous Ranked Probability Score (CRPS). This includes implementing CRPS-based losses and optimization procedures, aggregation or weighting strategies for ensemble members, and methods to produce calibrated, sharp probabilistic forecasts and to scale such training to larger ensembles.
This study addresses the challenges of derivative-dependent fitting, difficult calibration, and insufficient sharpness in stochastic parameterization schemes for ensemble forecasting. We propose an end-to-end trajectory learning framework directly optimized via the Continuous Ranked Probability Score (CRPS). Our method jointly models additive and multiplicative stochastic parameterizations using CRPS as the loss function—bypassing explicit derivative fitting—and integrates deep learning with ensemble forecasting within the two-scale Lorenz ’96 system. Our key contributions are: (i) the first application of CRPS-driven trajectory learning to stochastic parameterization training, simultaneously improving forecast accuracy and probabilistic sharpness; and (ii) a naturally well-calibrated model that seamlessly interfaces with data assimilation systems. Experiments demonstrate that our approach significantly outperforms conventional derivative-fitting methods in short-term ensemble forecasting, achieving superior accuracy, sharpness, and generalization across diverse dynamical regimes.
To address the slow sampling speed and high computational cost of diffusion models in regional ensemble weather forecasting, this paper proposes a single-forward-pass generative method trained directly on the Continuous Ranked Probability Score (CRPS). Innovatively integrating CRPS optimization into the Limited-Area Modeling (LAM) framework, the approach employs a single latent-variable sampling strategy coupled with conditional injection, enabling high-fidelity ensemble generation in one forward pass. Evaluated on the MEPS dataset, it achieves a 39× speedup over diffusion-based methods while preserving competitive forecast accuracy and fine-grained spatial structure. The core contribution is the first end-to-end ensemble generation framework that jointly optimizes for probabilistic calibration (via CRPS), computational efficiency (single-forward inference), and physical consistency within LAM—establishing a novel paradigm for efficient regional probabilistic forecasting.
Probabilistic weather forecasting faces a fundamental trade-off between preserving fine-scale structural fidelity and maintaining large-scale forecast skill. Method: To address this, we propose a multi-scale weighted adaptive forecast Continuous Ranked Probability Score (afCRPS) loss function—the first end-to-end scale-aware optimization of afCRPS. Our approach decomposes raw forecast fields into multi-scale components and applies scale-dependent weighting, thereby strengthening constraints on small-scale variability while preserving the differentiability of the overall CRPS. The framework is integrated into the AIFS-CRPS model. Results: Evaluated on ECMWF data, it significantly improves fine-scale structural fidelity for high-resolution fields—e.g., precipitation edge sharpness and local intensity—without degrading large-scale circulation or temperature forecast skill (as measured by ACC and RMSE). This work establishes a generalizable training paradigm for scale-adaptive probabilistic forecasting.
This study investigates the sensitivity of machine learning–based weather forecasting models to scale-aware scoring rules and their performance variations across different regions and spatial scales. Building upon the AIFS-CRPS framework, we present the first systematic evaluation of multivariate scale-aware loss functions—including the fair Continuous Ranked Probability Score (CRPS), global energy score, and graph energy score—for global probabilistic forecasting. Through spectral analysis, we elucidate how these losses influence the spectral structure of forecast fields. Our experiments demonstrate that explicitly incorporating scale-aware losses significantly enhances the realism of predicted atmospheric fields. Notably, the graph energy score yields optimal performance in tropical regions, whereas the global energy score exhibits slight degradation, thereby confirming the efficacy and potential of multivariate scoring rules in advancing machine learning–driven weather prediction.
Numerical weather prediction (NWP) ensembles often exhibit systematic biases and insufficient spread in extreme wind speed forecasting. To address this, we propose a novel EMOS statistical post-processing paradigm based on the threshold-weighted continuous ranked probability score (twCRPS). We derive closed-form expressions for twCRPS under multiple parametric distributions—first such derivation—and design a joint framework integrating weighted parameter estimation and linear pooling to simultaneously improve bulk calibration and tail sensitivity. Experiments on ECMWF ensemble forecasts demonstrate that our method significantly enhances probabilistic forecasting performance for extreme wind events across multiple thresholds. The approach balances theoretical rigor with operational practicality, offering a generalizable, threshold-aware training objective for extreme-weather probability forecasting.
This study addresses the challenge of learning predictive distributions in long-lead probabilistic weather forecasting, where high uncertainty complicates accurate modeling. To this end, it proposes a multi-noise-level framework based on distributed diffusion models, introducing an auxiliary conditional denoising task that leverages partial future information to reduce ambiguity. Furthermore, standard Continuous Ranked Probability Score (CRPS) training is extended to multiple noise levels, enabling the optimization of proper scoring rules through a single stochastic predictor by merely incorporating additional conditional inputs. The proposed approach significantly enhances the accuracy and calibration of global, high-dimensional weather forecasts while maintaining the computational efficiency of single-pass forward inference. Additionally, it demonstrates improved generalization capabilities under distribution shifts.
This study addresses the common trade-off in post-processing ensemble forecasts, where neural network–based calibration often sacrifices predictive sharpness—particularly for short lead times—to improve reliability. To jointly optimize both calibration and sharpness, the authors propose a novel approach that introduces, for the first time, a sharpness penalty term directly into the continuous ranked probability score (CRPS) loss function. Assuming a Gaussian predictive distribution, the method is evaluated on ECMWF 2-meter temperature ensemble forecasts and achieves a reduction of 8.2%–12.5% in the width of central prediction intervals while maintaining CRPS and mean RMSE at levels comparable to baseline methods. This demonstrates a significant enhancement in forecast sharpness without compromising overall forecasting skill.
Current evaluation methods for probabilistic forecasts rely on single scalar metrics, which fail to reveal the trade-off between sharpness and accuracy. This work proposes the Interval Score ROC curve (IS-ROC), the first geometric representation that fully characterizes families of interval forecasts across varying levels of sharpness while guaranteeing Pareto optimality and convexity. Leveraging the geometric properties of the IS-ROC, the authors further develop a tangent-based optimization method for calibration and a convex hull ensemble strategy. Experimental results demonstrate that this framework significantly outperforms existing approaches in terms of evaluation comprehensiveness, calibration quality, and ensemble performance.
This study addresses the lack of systematic evaluation of uncertainty reliability in probabilistic forecasting for physical systems, particularly between generative models and ensemble methods trained with the Continuous Ranked Probability Score (CRPS). The authors establish a unified evaluation framework to conduct the first systematic comparison of these two approaches under identical model scales and computational budgets in two-dimensional spatiotemporal systems. Results demonstrate that CRPS-based ensembles consistently achieve more reliable uncertainty coverage and faster inference in both single-step and roll-out predictions. Generative models exhibit comparable coverage only when trained in the original data space but incur significantly higher latency. Both approaches yield similar predictive accuracy. Notably, CRPS ensembles maintain strong performance even in latent spaces, highlighting their advantage in enabling efficient and reliable uncertainty quantification.
研究通过在推理时对确定性机器学习天气模型的权重进行随机扰动,以低成本生成不确定性估计,提高预报准确性。