Score
Designs and implements evaluation procedures and pipelines to score, compare, and diagnose probabilistic forecasts using proper scoring rules (e.g., Brier, energy), calibration and sharpness checks, and aggregated summary statistics; and produces comparative analyses across times, locations, or forecast horizons including baseline comparisons and retrospective performance assessments.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
本文评估了不同奖励函数对大语言模型预测性能和行为的影响,比较了五种适当评分规则作为训练目标的效果。
This paper addresses the limitation of conventional weather forecast evaluation—its overreliance on statistical accuracy while neglecting decision utility—by proposing a novel value-oriented evaluation paradigm grounded in the decision-maker’s perspective. Methodologically, it introduces a “decision calibration” framework that integrates decision theory, probabilistic calibration analysis, and multi-task utility assessment to systematically compare machine learning and numerical weather prediction models under realistic decision-making scenarios. The key contribution is the empirical revelation of a substantial mismatch between statistical performance and decision utility: the same forecast model exhibits markedly divergent rankings across distinct decision tasks (e.g., disaster mitigation vs. energy dispatch), rendering traditional metrics inadequate for application-specific model selection. The framework establishes an interpretable, task-adapted foundation for quantifying forecast service value and guiding operational model choice.
This study addresses the limited interpretability of the Brier score in diagnosing deficiencies in probabilistic forecasts by proposing an algebraic rearrangement based on Yates’ covariance decomposition. The method cleanly decomposes the Brier score into three non-negative components: variance mismatch, insufficient correlation, and overall calibration bias. This decomposition is not only mathematically concise but also highly interpretable, explicitly revealing that perfect prediction requires simultaneous satisfaction of three conditions: matched variances, perfect positive correlation, and agreement in means. By elucidating the distinct sources of forecast error, the approach substantially enhances the diagnostic capability for evaluating probabilistic predictions and provides both a theoretical foundation and a practical tool for improving predictive models.
Existing theoretical characterizations of proper scoring rules for probabilistic forecasting and distribution estimation are fragmented and lack methodological clarity. Method: Drawing on convex analysis, information geometry, and decision theory, this paper systematically unifies general characterization theorems with canonical rule families—including logarithmic score and Brier score—establishing rigorous criteria for score propriety and a principled optimization framework. Contribution/Results: We prove, for the first time, the equivalence between proper scoring rules as unbiased estimation tools and as consistent evaluation criteria for probabilistic forecasts. This foundational result provides a unified theoretical basis for Bayesian updating, density estimation, and model calibration. It significantly extends the methodological scope and applicability of proper scoring rules in statistical inference and machine learning, clarifying their role in both theoretical foundations and practical algorithm design.
This study addresses the limitation of existing recalibration methods, which often obscure miscalibration in specific regions such as extreme events. To overcome this, we propose an outcome-conditional recalibration post-processing method that leverages quantile recalibration and conditional distribution scaling to achieve precise correction of arbitrary predictive distributions within user-defined regions. By simultaneously preserving global performance and local reliability, the proposed approach significantly enhances conditional calibration on regression benchmark tasks. Furthermore, when applied to electricity price forecasting, it substantially improves calibration in negative-price regimes with negligible accuracy loss. Overall, this work provides more reliable localized guarantees for probabilistic forecasting.
This study addresses a critical limitation in existing probabilistic electricity price forecasting methods, which overly prioritize sharpness at the expense of calibration, yielding overconfident and statistically unreliable uncertainty estimates. The authors systematically analyze the trade-off between calibration and sharpness, demonstrating how prevailing scoring rules—by neglecting reliability—distort predictive distributions and risk degenerating probabilistic models into mere surrogates of deterministic forecasts. To remedy this, the paper proposes a theoretical framework that elevates calibration to a central modeling principle, integrating probabilistic prediction, calibration assessment, and proper scoring rules. It advocates for the development of calibration-aware predictive objectives and architectures, offering a principled direction to enhance the reliability and comprehensiveness of forecasts in energy markets.
该研究解决了排名对比中的虚假关联问题,通过固定学习映射、参考律和系数行总和,并使用无偏三轨迹核估计交互作用来认证真实预测技能。
Current evaluation methods for probabilistic forecasts rely on single scalar metrics, which fail to reveal the trade-off between sharpness and accuracy. This work proposes the Interval Score ROC curve (IS-ROC), the first geometric representation that fully characterizes families of interval forecasts across varying levels of sharpness while guaranteeing Pareto optimality and convexity. Leveraging the geometric properties of the IS-ROC, the authors further develop a tangent-based optimization method for calibration and a convex hull ensemble strategy. Experimental results demonstrate that this framework significantly outperforms existing approaches in terms of evaluation comprehensiveness, calibration quality, and ensemble performance.
This study addresses the limitation of existing forecasting systems that rely predominantly on point predictions and thus fail to adequately characterize uncertainty for informed decision-making. To overcome this, the authors propose a hybrid framework that extends point forecasts from classical models—such as Theta, exponential smoothing, and ARIMA—into probabilistic forecasts by integrating error post-processing with model-specific, horizon-dependent uncertainty scaling. The approach calibrates forecast errors using historical simulation, conformal prediction, quantile regression, and GARCH-based methods, and systematically evaluates in-sample versus out-of-sample calibration performance. Empirical results on the M4 dataset demonstrate an average 4.6% reduction in Continuous Ranked Probability Score (CRPS). In-sample calibration consistently outperforms out-of-sample calibration, particularly over longer forecast horizons, thereby validating the effectiveness and practical utility of the proposed framework.