Score
Design and train probabilistic predictive models and the corresponding loss functions using proper scoring rules (for example energy score or other strictly proper scoring objectives) so that optimization drives model outputs toward calibrated, accurate predictive distributions. Implement training procedures, sampling or gradient-estimation methods, and evaluation metrics that directly optimize or assess probabilistic accuracy across predicted distributions.
This paper addresses the design of proper scoring rules for multidimensional forecasting settings, aiming to incentivize forecasters to exert effort and truthfully report their beliefs. Methodologically, it introduces the first optimization framework explicitly targeting *effort incentives*, integrating game-theoretic modeling with convex optimization. For simple settings, it derives closed-form characterizations of optimal rules; for general cases, it develops an efficient and exact algorithm; and it identifies several structurally simple approximate rules with near-optimal performance. Theoretical analysis reveals that classical proper scoring rules—such as the quadratic score—can substantially deviate from optimality under multidimensional effort. In contrast, the proposed algorithm computes exact optimal rules, while the simple approximations achieve over 95% of the optimal incentive efficiency. These results establish a new paradigm for information design and prediction market mechanisms, bridging incentive alignment with practical implementability.
Existing theoretical characterizations of proper scoring rules for probabilistic forecasting and distribution estimation are fragmented and lack methodological clarity. Method: Drawing on convex analysis, information geometry, and decision theory, this paper systematically unifies general characterization theorems with canonical rule families—including logarithmic score and Brier score—establishing rigorous criteria for score propriety and a principled optimization framework. Contribution/Results: We prove, for the first time, the equivalence between proper scoring rules as unbiased estimation tools and as consistent evaluation criteria for probabilistic forecasts. This foundational result provides a unified theoretical basis for Bayesian updating, density estimation, and model calibration. It significantly extends the methodological scope and applicability of proper scoring rules in statistical inference and machine learning, clarifying their role in both theoretical foundations and practical algorithm design.
Traditional probabilistic models are typically trained using task-agnostic log-loss, which often yields propensity scores with large errors, high bias, and high variance in boundary regions—particularly detrimental in causal inference tasks such as inverse probability weighting. This work proposes a general framework that, for the first time, integrates the error structure of downstream tasks into the design of strictly proper scoring rules. By aligning the local curvature of the scoring rule with that of the target loss, the authors derive a closed-form loss function tailored for average treatment effect estimation, along with its associated canonical probability mapping, enabling end-to-end task-oriented training. The approach is compatible with both neural networks and gradient boosting models and consistently outperforms standard log-likelihood and covariate balancing methods across multiple causal inference benchmarks, substantially improving estimation accuracy and stability.
本文评估了不同奖励函数对大语言模型预测性能和行为的影响,比较了五种适当评分规则作为训练目标的效果。
This paper challenges the unverified implicit assumption in the predict-then-optimize paradigm that “higher prediction accuracy necessarily yields better downstream decisions,” particularly in multiclass classification settings. Method: We propose a controllable, interpretable multiclass prediction simulation framework that explicitly models error types and distributions, enabling systematic analysis of how classification errors affect decision quality in constrained optimization. Contribution/Results: Experiments on job scheduling and other combinatorial optimization tasks reveal a nonlinear relationship between prediction error and decision performance: improving prediction accuracy does not guarantee improved solution quality—and can even degrade decisions when error patterns shift. Our findings question the conventional coupling logic between prediction and optimization, providing theoretical foundations and practical guidance for designing, evaluating, and calibrating classifiers specifically tailored to decision objectives.
Existing generative models struggle to align prediction errors with downstream decision costs in high-stakes scenarios. This work proposes a decision-aware training approach that explicitly embeds a differentiable decision loss into the energy score objective, establishing a joint optimization framework grounded in proper scoring rules. By directly penalizing prediction biases that incur high decision costs while preserving full probabilistic forecasts, the method achieves both theoretical rigor and practical relevance. Empirical evaluations on synthetic data and two real-world tasks demonstrate substantial improvements in predictive performance within cost-sensitive regions, without compromising the overall quality of the probabilistic outputs.
This work proposes a prediction-oriented Bayesian inference approach under model misspecification, which constructs a posterior distribution that balances predictive performance and uncertainty quantification by optimizing a scoring rule—such as the logarithmic score—over predictive distributions, augmented with a φ-divergence regularizer relative to a reference prior. Leveraging a finite-dimensional dual formulation, the method establishes theoretical optimality under a zero duality gap condition and derives finite-sample bounds on predictive risk for the resulting approximate posterior. By employing a semi-analytical posterior representation and solving the associated dual optimization problem, the approach demonstrates superior predictive accuracy and numerical stability in both classification tasks and experiments involving Gaussian mixture misspecification.
This study addresses a critical limitation in existing probabilistic electricity price forecasting methods, which overly prioritize sharpness at the expense of calibration, yielding overconfident and statistically unreliable uncertainty estimates. The authors systematically analyze the trade-off between calibration and sharpness, demonstrating how prevailing scoring rules—by neglecting reliability—distort predictive distributions and risk degenerating probabilistic models into mere surrogates of deterministic forecasts. To remedy this, the paper proposes a theoretical framework that elevates calibration to a central modeling principle, integrating probabilistic prediction, calibration assessment, and proper scoring rules. It advocates for the development of calibration-aware predictive objectives and architectures, offering a principled direction to enhance the reliability and comprehensiveness of forecasts in energy markets.
This study investigates the sensitivity of machine learning–based weather forecasting models to scale-aware scoring rules and their performance variations across different regions and spatial scales. Building upon the AIFS-CRPS framework, we present the first systematic evaluation of multivariate scale-aware loss functions—including the fair Continuous Ranked Probability Score (CRPS), global energy score, and graph energy score—for global probabilistic forecasting. Through spectral analysis, we elucidate how these losses influence the spectral structure of forecast fields. Our experiments demonstrate that explicitly incorporating scale-aware losses significantly enhances the realism of predicted atmospheric fields. Notably, the graph energy score yields optimal performance in tropical regions, whereas the global energy score exhibits slight degradation, thereby confirming the efficacy and potential of multivariate scoring rules in advancing machine learning–driven weather prediction.
This study addresses the challenge of learning predictive distributions in long-lead probabilistic weather forecasting, where high uncertainty complicates accurate modeling. To this end, it proposes a multi-noise-level framework based on distributed diffusion models, introducing an auxiliary conditional denoising task that leverages partial future information to reduce ambiguity. Furthermore, standard Continuous Ranked Probability Score (CRPS) training is extended to multiple noise levels, enabling the optimization of proper scoring rules through a single stochastic predictor by merely incorporating additional conditional inputs. The proposed approach significantly enhances the accuracy and calibration of global, high-dimensional weather forecasts while maintaining the computational efficiency of single-pass forward inference. Additionally, it demonstrates improved generalization capabilities under distribution shifts.