Score
Designs and evaluates models and systems that predict future values, states, or events over time (including point and probabilistic forecasts and selection of forecast horizons). Builds forecasting pipelines and experiments, performs backtesting and comparison of algorithms, and analyzes forecast accuracy, calibration, and uncertainty to inform decision-making.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This paper addresses the limitation of conventional weather forecast evaluation—its overreliance on statistical accuracy while neglecting decision utility—by proposing a novel value-oriented evaluation paradigm grounded in the decision-maker’s perspective. Methodologically, it introduces a “decision calibration” framework that integrates decision theory, probabilistic calibration analysis, and multi-task utility assessment to systematically compare machine learning and numerical weather prediction models under realistic decision-making scenarios. The key contribution is the empirical revelation of a substantial mismatch between statistical performance and decision utility: the same forecast model exhibits markedly divergent rankings across distinct decision tasks (e.g., disaster mitigation vs. energy dispatch), rendering traditional metrics inadequate for application-specific model selection. The framework establishes an interpretable, task-adapted foundation for quantifying forecast service value and guiding operational model choice.
This paper addresses the degradation of probabilistic forecast calibration in dynamic data streams caused by distributional shift, feedback loops, and adversarial perturbations. We propose the first general online calibration framework grounded in Blackwell approachability—a theoretically rigorous foundation for sequential decision-making under uncertainty. Our method provides strong calibration guarantees in compact output spaces (e.g., classification and bounded regression) and enables lossless post-hoc recalibration of arbitrary pre-trained predictors. Technically, it unifies insights from Blackwell approachability theory, online optimization, and gradient-based updates, and introduces task-specific efficient algorithms for both classification and regression. Empirical evaluation demonstrates substantial improvements in calibration quality for energy system forecasting, with marked gains in robustness and practical utility for downstream decision-making tasks.
This work addresses two fundamental questions: “What constitutes a principled predictive evaluation metric?” and “How are existing metrics formally related?” We propose a unified evaluation framework grounded in game theory, introducing the first four-dimensional “predictive welfare” metric—comprising calibration, predictability, randomness, and regret—to holistically assess prediction quality. Theoretically, we rigorously prove the equivalence between calibration and regret, and establish a duality between predictive superiority and outcome randomness. Methodologically, we formalize the framework by integrating probabilistic calibration, regret analysis, and algorithmic randomness measures—specifically Martingale difference sequences. Our approach provides a more rigorous theoretical foundation for predictive evaluation, significantly enhancing both the interpretability and robustness assessment of trustworthy AI predictions.
Existing probabilistic forecasting evaluation methods lack the ability to characterize tail calibration—critical for high-impact extreme events, whose reliability is increasingly vital for risk-informed decision-making. Method: This paper introduces, for the first time, a general definition of tail calibration, rigorously connecting it to classical probabilistic calibration theory and integrating the Peaks-over-Threshold (POT) framework from extreme value theory. We develop an operational diagnostic framework by unifying probabilistic calibration theory, extreme-value statistics, diagnostic statistical tests, and empirical analysis. Contribution/Results: Applied to European precipitation forecasts, our framework significantly improves the quantification of predictive credibility for high-impact, rare events. It enables rigorous assessment of tail behavior in probabilistic forecasts and establishes a novel paradigm for extreme-event risk assessment and decision support.
This study addresses the limitation of existing forecasting systems that rely predominantly on point predictions and thus fail to adequately characterize uncertainty for informed decision-making. To overcome this, the authors propose a hybrid framework that extends point forecasts from classical models—such as Theta, exponential smoothing, and ARIMA—into probabilistic forecasts by integrating error post-processing with model-specific, horizon-dependent uncertainty scaling. The approach calibrates forecast errors using historical simulation, conformal prediction, quantile regression, and GARCH-based methods, and systematically evaluates in-sample versus out-of-sample calibration performance. Empirical results on the M4 dataset demonstrate an average 4.6% reduction in Continuous Ranked Probability Score (CRPS). In-sample calibration consistently outperforms out-of-sample calibration, particularly over longer forecast horizons, thereby validating the effectiveness and practical utility of the proposed framework.
This study addresses the challenge of uncertainty quantification in aggregated time series forecasting, particularly for annual totals and year-over-year growth rates. It proposes a simulation-augmented multi-step split conformal prediction method (SA-MSCP), which generates future trajectories via block bootstrap resampling from cross-validated residuals and constructs calibrated prediction intervals using empirical quantiles. By innovatively integrating a simulation-augmentation mechanism into the multi-step split conformal prediction framework, the method significantly improves empirical coverage for both aggregate totals and their growth rates, yielding more reliable uncertainty estimates without compromising predictive accuracy.
This work proposes a machine learning–oriented paradigm for weather forecasting that reimagines the traditionally complex and closed operational systems to meet the demands of efficiency, openness, and collaboration in the era of artificial intelligence. By integrating agent-driven software engineering, open compressed data formats, shared validation workflows, interactive computing environments, and generative AI techniques, the framework systematically transforms model development, data utilization, computational management, and service delivery. Designed to equip meteorological and climate centers with future-ready infrastructure, it establishes robust data governance mechanisms, quality assurance protocols, and pathways for workforce skill transformation. The approach maintains scientific rigor while substantially enhancing the accessibility, efficiency, and interactivity of forecasting services.
This work addresses the challenge that existing calibration tests for conditional quantile predictors struggle to handle distributional shifts and discrepancies in information sets, lacking feature-aware, continuous monitoring capabilities. The authors propose a distribution-free, game-theoretic sequential auditing framework that formally defines conditional quantile calibration under varying feature information sets—a notion not previously established—and provides finite-time detection guarantees without requiring independent and identically distributed data. By integrating contextual linear betting strategies with nonparametric e-processes, the method enables interpretable, feature-level calibration audits. Empirical evaluations demonstrate that the framework effectively detects significant miscalibration in state-of-the-art time series models, such as Chronos-2, across critical features.