proper scoring rules

Designing and applying scoring rules and aggregation methods that produce calibrated, sharp probabilistic and uncertainty estimates and enable principled evaluation and conversion between predictor types. Used to preserve sample complexity in omnipredictor conversions, train long-horizon covariance-aware losses, and optimize out-of-sample predictive performance.

properscoringrules

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Proper scoring rules for estimation and forecast evaluation

Apr 02, 2025
KG
Kartik G. Waghmare
🏛️ ETH Zurich

Existing theoretical characterizations of proper scoring rules for probabilistic forecasting and distribution estimation are fragmented and lack methodological clarity. Method: Drawing on convex analysis, information geometry, and decision theory, this paper systematically unifies general characterization theorems with canonical rule families—including logarithmic score and Brier score—establishing rigorous criteria for score propriety and a principled optimization framework. Contribution/Results: We prove, for the first time, the equivalence between proper scoring rules as unbiased estimation tools and as consistent evaluation criteria for probabilistic forecasts. This foundational result provides a unified theoretical basis for Bayesian updating, density estimation, and model calibration. It significantly extends the methodological scope and applicability of proper scoring rules in statistical inference and machine learning, clarifying their role in both theoretical foundations and practical algorithm design.

Applying scoring rules to probability distribution estimationCharacterizing proper scoring rules mathematicallyEvaluating forecasts using proper scoring rules

Sample-Efficient Omniprediction for Proper Losses

Oct 14, 2025
IG
Isaac Gibbs
🏛️ University of California, Berkeley

This work addresses the problem of designing a single probabilistic predictor that enables multiple downstream decision-makers—each operating under a distinct, admissible loss function—to achieve optimal decisions. Existing approaches suffer from reliance on randomization, high sample complexity, and difficulty in simultaneously optimizing for multiple losses. To overcome these limitations, we propose a structure-aware deterministic algorithm: leveraging the convex conjugate structure of proper losses and an adversarial game framework based on online-to-batch conversion, our method jointly optimizes over multicalibration constraints and the given set of loss functions. The resulting predictor is deterministic, avoids randomized forecasts, and achieves significantly lower sample complexity. Theoretically grounded, it surpasses prior boosting-based methods and—crucially—breaks the long-standing sample-size bottleneck imposed by multicalibration requirements. Our approach delivers a simple, efficient, and provably optimal solution for multi-loss-compatible prediction.

Constructing probabilistic predictions for accurate downstream decision-makingDesigning single predictors minimizing multiple proper losses simultaneouslyDeveloping sample-efficient omniprediction algorithms without complex randomization

Optimization of Scoring Rules

Jul 06, 2020
YL
Yingkai Li
🏛️ Yale University | Northwestern University | Toyota Technological Institute at Chicago

This paper addresses the design of proper scoring rules for multidimensional forecasting settings, aiming to incentivize forecasters to exert effort and truthfully report their beliefs. Methodologically, it introduces the first optimization framework explicitly targeting *effort incentives*, integrating game-theoretic modeling with convex optimization. For simple settings, it derives closed-form characterizations of optimal rules; for general cases, it develops an efficient and exact algorithm; and it identifies several structurally simple approximate rules with near-optimal performance. Theoretical analysis reveals that classical proper scoring rules—such as the quadratic score—can substantially deviate from optimality under multidimensional effort. In contrast, the proposed algorithm computes exact optimal rules, while the simple approximations achieve over 95% of the optimal incentive efficiency. These results establish a new paradigm for information design and prediction market mechanisms, bridging incentive alignment with practical implementability.

Comparing optimal scoring rules with standard alternativesDesigning incentives for multi-dimensional information acquisitionOptimizing scoring rules for truthful information reporting

This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.

calibrationclassificationhierarchical relations

Calibrated Probabilistic Forecasts for Arbitrary Sequences

Sep 27, 2024
CM
Charles Marx
🏛️ Stanford University | Cornell Tech

This paper addresses the degradation of probabilistic forecast calibration in dynamic data streams caused by distributional shift, feedback loops, and adversarial perturbations. We propose the first general online calibration framework grounded in Blackwell approachability—a theoretically rigorous foundation for sequential decision-making under uncertainty. Our method provides strong calibration guarantees in compact output spaces (e.g., classification and bounded regression) and enables lossless post-hoc recalibration of arbitrary pre-trained predictors. Technically, it unifies insights from Blackwell approachability theory, online optimization, and gradient-based updates, and introduces task-specific efficient algorithms for both classification and regression. Empirical evaluation demonstrates substantial improvements in calibration quality for energy system forecasting, with marked gains in robustness and practical utility for downstream decision-making tasks.

Ensures valid uncertainty estimates for evolving data streams.Guarantees calibrated uncertainties in compact outcome spaces.Recalibrates existing forecasters without losing predictive performance.

Latest Papers

What's happening recently
View more

This work formally introduces conditional risk calibration—the estimation of a predictive model’s expected loss given an input—as a standalone machine learning problem, framing it as a standard regression task. Through theoretical analysis, it uncovers an intrinsic connection between conditional risk calibration and individual probability calibration, offering a novel perspective on the Learn then Decide (L2D) framework. By integrating techniques from regression modeling, probability calibration, and uncertainty quantification, the authors conduct systematic experiments across both classification and regression settings. The results demonstrate the effectiveness of the proposed approach, highlighting its practical utility in uncertainty-aware decision-making through both qualitative insights and quantitative improvements.

conditional probabilityconditional riskexpected loss

This study investigates the sensitivity of machine learning–based weather forecasting models to scale-aware scoring rules and their performance variations across different regions and spatial scales. Building upon the AIFS-CRPS framework, we present the first systematic evaluation of multivariate scale-aware loss functions—including the fair Continuous Ranked Probability Score (CRPS), global energy score, and graph energy score—for global probabilistic forecasting. Through spectral analysis, we elucidate how these losses influence the spectral structure of forecast fields. Our experiments demonstrate that explicitly incorporating scale-aware losses significantly enhances the realism of predicted atmospheric fields. Notably, the graph energy score yields optimal performance in tropical regions, whereas the global energy score exhibits slight degradation, thereby confirming the efficacy and potential of multivariate scoring rules in advancing machine learning–driven weather prediction.

forecast verificationmachine learningprobabilistic weather forecasting

This work addresses the challenge of effectively aggregating statistical evidence under unknown dependence structures by proposing a unified framework grounded in permutation invariance. The approach constructs exchangeable data units, aggregates statistics within transformed datasets, and calibrates results across transformations, accommodating single-batch, sequential, and two-stage strategies. By integrating group invariance, exchangeability modeling, sequential alpha-spending, and a decoupling of standardization from calibration, the method achieves high power and adaptivity in finite samples, substantially outperforming traditional calibration techniques such as Bonferroni correction. Empirical evaluations demonstrate that the framework guarantees valid inference under arbitrary dependence structures in tasks including nonparametric testing and conformal prediction, while supporting data-driven aggregation rules and early rejection mechanisms.

calibrationexchangeabilitypermutation-based inference

This study addresses a critical limitation in existing probabilistic electricity price forecasting methods, which overly prioritize sharpness at the expense of calibration, yielding overconfident and statistically unreliable uncertainty estimates. The authors systematically analyze the trade-off between calibration and sharpness, demonstrating how prevailing scoring rules—by neglecting reliability—distort predictive distributions and risk degenerating probabilistic models into mere surrogates of deterministic forecasts. To remedy this, the paper proposes a theoretical framework that elevates calibration to a central modeling principle, integrating probabilistic prediction, calibration assessment, and proper scoring rules. It advocates for the development of calibration-aware predictive objectives and architectures, offering a principled direction to enhance the reliability and comprehensiveness of forecasts in energy markets.

calibrationelectricity priceprobabilistic forecasting

This work resolves the open question of whether randomization is necessary to achieve optimal sample complexity in multi-calibration. The authors present the first deterministic algorithm that attains ε-multi-calibration for any collection of group weight functions, and naturally extends to outcome indistinguishability (OI) and omniprediction. Built upon a deterministic learning framework coupled with a covering-based test set technique, the method achieves the minimax-optimal sample complexity of $\tilde{O}(\varepsilon^{-3})$, thereby establishing for the first time that randomization is not essential. This result simultaneously settles several long-standing open problems across multi-calibration, OI, and omniprediction in a unified manner.

deterministic predictormulticalibrationomniprediction

Hot Scholars

SK

Sanmi Koyejo

Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics
MT

Mike Thelwall

School of Information, Journalism and Communication, The University of Sheffield
scientometricsaltmetricssentiment analysissocial media
XZ

Xiaoming Zhai

Associate Professor, University of Georgia
Science EducationAIAssessment
HJ

Hong Jiao

University of Maryland, College Park
educational measurementpsychometrics
SC

Shuheng Chen

University of Southern California
Machine LearningData SciencePredictive AnalyticsClinical Prediction