Score
Designing and applying scoring rules and aggregation methods that produce calibrated, sharp probabilistic and uncertainty estimates and enable principled evaluation and conversion between predictor types. Used to preserve sample complexity in omnipredictor conversions, train long-horizon covariance-aware losses, and optimize out-of-sample predictive performance.
Existing theoretical characterizations of proper scoring rules for probabilistic forecasting and distribution estimation are fragmented and lack methodological clarity. Method: Drawing on convex analysis, information geometry, and decision theory, this paper systematically unifies general characterization theorems with canonical rule families—including logarithmic score and Brier score—establishing rigorous criteria for score propriety and a principled optimization framework. Contribution/Results: We prove, for the first time, the equivalence between proper scoring rules as unbiased estimation tools and as consistent evaluation criteria for probabilistic forecasts. This foundational result provides a unified theoretical basis for Bayesian updating, density estimation, and model calibration. It significantly extends the methodological scope and applicability of proper scoring rules in statistical inference and machine learning, clarifying their role in both theoretical foundations and practical algorithm design.
This work addresses the problem of designing a single probabilistic predictor that enables multiple downstream decision-makers—each operating under a distinct, admissible loss function—to achieve optimal decisions. Existing approaches suffer from reliance on randomization, high sample complexity, and difficulty in simultaneously optimizing for multiple losses. To overcome these limitations, we propose a structure-aware deterministic algorithm: leveraging the convex conjugate structure of proper losses and an adversarial game framework based on online-to-batch conversion, our method jointly optimizes over multicalibration constraints and the given set of loss functions. The resulting predictor is deterministic, avoids randomized forecasts, and achieves significantly lower sample complexity. Theoretically grounded, it surpasses prior boosting-based methods and—crucially—breaks the long-standing sample-size bottleneck imposed by multicalibration requirements. Our approach delivers a simple, efficient, and provably optimal solution for multi-loss-compatible prediction.
This paper addresses the design of proper scoring rules for multidimensional forecasting settings, aiming to incentivize forecasters to exert effort and truthfully report their beliefs. Methodologically, it introduces the first optimization framework explicitly targeting *effort incentives*, integrating game-theoretic modeling with convex optimization. For simple settings, it derives closed-form characterizations of optimal rules; for general cases, it develops an efficient and exact algorithm; and it identifies several structurally simple approximate rules with near-optimal performance. Theoretical analysis reveals that classical proper scoring rules—such as the quadratic score—can substantially deviate from optimality under multidimensional effort. In contrast, the proposed algorithm computes exact optimal rules, while the simple approximations achieve over 95% of the optimal incentive efficiency. These results establish a new paradigm for information design and prediction market mechanisms, bridging incentive alignment with practical implementability.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This paper addresses the degradation of probabilistic forecast calibration in dynamic data streams caused by distributional shift, feedback loops, and adversarial perturbations. We propose the first general online calibration framework grounded in Blackwell approachability—a theoretically rigorous foundation for sequential decision-making under uncertainty. Our method provides strong calibration guarantees in compact output spaces (e.g., classification and bounded regression) and enables lossless post-hoc recalibration of arbitrary pre-trained predictors. Technically, it unifies insights from Blackwell approachability theory, online optimization, and gradient-based updates, and introduces task-specific efficient algorithms for both classification and regression. Empirical evaluation demonstrates substantial improvements in calibration quality for energy system forecasting, with marked gains in robustness and practical utility for downstream decision-making tasks.
This work formally introduces conditional risk calibration—the estimation of a predictive model’s expected loss given an input—as a standalone machine learning problem, framing it as a standard regression task. Through theoretical analysis, it uncovers an intrinsic connection between conditional risk calibration and individual probability calibration, offering a novel perspective on the Learn then Decide (L2D) framework. By integrating techniques from regression modeling, probability calibration, and uncertainty quantification, the authors conduct systematic experiments across both classification and regression settings. The results demonstrate the effectiveness of the proposed approach, highlighting its practical utility in uncertainty-aware decision-making through both qualitative insights and quantitative improvements.
This study investigates the sensitivity of machine learning–based weather forecasting models to scale-aware scoring rules and their performance variations across different regions and spatial scales. Building upon the AIFS-CRPS framework, we present the first systematic evaluation of multivariate scale-aware loss functions—including the fair Continuous Ranked Probability Score (CRPS), global energy score, and graph energy score—for global probabilistic forecasting. Through spectral analysis, we elucidate how these losses influence the spectral structure of forecast fields. Our experiments demonstrate that explicitly incorporating scale-aware losses significantly enhances the realism of predicted atmospheric fields. Notably, the graph energy score yields optimal performance in tropical regions, whereas the global energy score exhibits slight degradation, thereby confirming the efficacy and potential of multivariate scoring rules in advancing machine learning–driven weather prediction.
This work addresses the challenge of effectively aggregating statistical evidence under unknown dependence structures by proposing a unified framework grounded in permutation invariance. The approach constructs exchangeable data units, aggregates statistics within transformed datasets, and calibrates results across transformations, accommodating single-batch, sequential, and two-stage strategies. By integrating group invariance, exchangeability modeling, sequential alpha-spending, and a decoupling of standardization from calibration, the method achieves high power and adaptivity in finite samples, substantially outperforming traditional calibration techniques such as Bonferroni correction. Empirical evaluations demonstrate that the framework guarantees valid inference under arbitrary dependence structures in tasks including nonparametric testing and conformal prediction, while supporting data-driven aggregation rules and early rejection mechanisms.
This study addresses a critical limitation in existing probabilistic electricity price forecasting methods, which overly prioritize sharpness at the expense of calibration, yielding overconfident and statistically unreliable uncertainty estimates. The authors systematically analyze the trade-off between calibration and sharpness, demonstrating how prevailing scoring rules—by neglecting reliability—distort predictive distributions and risk degenerating probabilistic models into mere surrogates of deterministic forecasts. To remedy this, the paper proposes a theoretical framework that elevates calibration to a central modeling principle, integrating probabilistic prediction, calibration assessment, and proper scoring rules. It advocates for the development of calibration-aware predictive objectives and architectures, offering a principled direction to enhance the reliability and comprehensiveness of forecasts in energy markets.
This work resolves the open question of whether randomization is necessary to achieve optimal sample complexity in multi-calibration. The authors present the first deterministic algorithm that attains ε-multi-calibration for any collection of group weight functions, and naturally extends to outcome indistinguishability (OI) and omniprediction. Built upon a deterministic learning framework coupled with a covering-based test set technique, the method achieves the minimax-optimal sample complexity of $\tilde{O}(\varepsilon^{-3})$, thereby establishing for the first time that randomization is not essential. This result simultaneously settles several long-standing open problems across multi-calibration, OI, and omniprediction in a unified manner.