Score
Design and implement algorithms and systems that fuse, calibrate, and evaluate heterogeneous model or expert outputs — scalar scores, probability distributions, intervals, confidences, or ranks — into unified quantitative assessments used for ranking, anomaly detection, decision thresholds, or training/reward signals. This includes building probabilistic scoring and calibration modules (log-probability, probabilistic scoring rules such as the continuous ranked probability score, interval and confidence-based scoring), aggregation strategies (majority-vote, weighted/ensemble, multimodal, multi-objective and rank-aware fusion), specialized score representations (e.g., quintuple score), and evaluation pipelines to compare and tune fusion methods.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This work addresses the limitations of human-centric evaluation—namely, the absence of verifiable ground truth, heterogeneity among expert judgments, and inconsistent rating scales—and the reliance of purely model-based evaluation on imperfect proxy metrics. To overcome these challenges, the authors propose AtC, a two-stage framework: first, it aggregates rankings by modeling annotator reliability to produce a consensus ranking; second, it calibrates arbitrary model scores onto this ranking via isotonic projection, preserving both ordinal consistency and quantitative information. This study is the first to integrate judgment aggregation with model-free calibration, providing theoretical guarantees of improved estimation efficiency, risk bounds under consensus misspecification, and asymptotic superiority over single-paradigm evaluation approaches. Experiments demonstrate that AtC significantly outperforms human-only or model-only methods on both semi-synthetic and real-world datasets.
This work proposes an interpretable statistical inference framework that decomposes predictive scoring functions into three components—calibration error, discrimination ability, and uncertainty—applicable to multi-step-ahead point forecasts such as means and quantiles, and compatible with both smooth and nonsmooth scoring rules. Building on linear recalibration and integrating Mincer–Zarnowitz regression with asymptotic inference theory, the method delivers the first fully interpretable tripartite decomposition for general scoring functions, unifying and extending classical calibration tests and predictive performance evaluation. Empirical applications to inflation surveys and financial risk models reveal critical discrepancies obscured by aggregate scores, exposing a misalignment between backtesting practices and predictive accuracy in banking regulation, thereby substantially enhancing the informativeness and statistical power of forecast evaluation.
Integrating hypothesis testing results across heterogeneous multi-source studies—some reporting only binary significance decisions, others only FDR control levels—poses a fundamental challenge for rigorous, unified FDR control. Method: We propose the Integrated Ranking and Thresholding (IRT) framework, which operates solely on binary rejection decisions, a prespecified global FDR level, and the set of hypotheses—requiring neither raw data, p-values, nor effect sizes. IRT employs nonparametric evidence aggregation and a ranking-driven thresholding mechanism, circumventing traditional meta-analysis assumptions of statistical homogeneity and reliance on shared summary statistics. Contribution/Results: IRT is the first method to achieve theoretically guaranteed strong FDR control under non-shared statistical summaries. We prove its FDR control property rigorously; simulations demonstrate superior performance over state-of-the-art integration methods; and real-world application to multi-center genome-wide association studies confirms its practical utility and robustness.
Probabilistic outputs of AI models often exhibit miscalibration—i.e., predicted confidence scores poorly reflect true accuracy—hindering their reliable deployment in safety-critical applications and ensemble systems. Method: This paper presents a systematic survey of probabilistic calibration evaluation methods for classification and object detection models. Grounded in statistical assessment theory, it unifies diverse approaches—including reliability diagrams, Brier score, expected calibration error (ECE), maximum calibration error (MCE), Kolmogorov–Smirnov test, ROC-based metrics, and IoU-aware measures—within a coherent framework covering binary, multiclass, and detection tasks. Contribution: We propose the first taxonomy of calibration metrics, categorizing 82 existing measures into four families: point-wise, binning-based, kernel/curve-based, and cumulative. Additionally, we introduce the first structured calibration metric knowledge base, enabling rapid metric selection, implementation, and comparative analysis—thereby establishing new interpretable and quantifiable benchmarks for trustworthy AI.
This work addresses the lack of explicit characterization of the interplay among information, reliability, and uncertainty in existing probabilistic forecast calibration methods. For any proper scoring rule, the authors propose the first general triple-decomposition framework grounded in information algebra and conditional entropy theory, which rigorously decomposes predictive loss into three distinct components: reliability (calibration error), information loss, and irreducible uncertainty. This framework uniquely quantifies the information loss incurred when mapping features to predictive scores and provides a unified interpretation of post-hoc calibration, model ensembling, and boosting strategies. In classification tasks, the approach is successfully applied to calibration evaluation, model aggregation, and staged training, clearly disentangling each component’s contribution to overall predictive uncertainty.
Existing methods struggle to effectively and interpretably evaluate and recalibrate the probabilistic outputs of black-box multiclass models without internal access. This work proposes the Multiclass Linear Log-Odds (MCLLO) recalibration framework, which, for the first time, enables calibration assessment and adjustment using only predicted probabilities from a single model. By modeling calibration through a linear transformation in log-odds space and employing a likelihood ratio test for direct calibration evaluation, MCLLO achieves both interpretability and broad applicability. Experiments across three real-world domains—image classification, obesity analysis, and ecological modeling—demonstrate that MCLLO matches or outperforms four state-of-the-art recalibration methods in terms of calibration performance.
This study addresses the longstanding disconnect between student academic performance prediction and metacognitive calibration by proposing a Unified Behavioral Prediction and Calibration Analysis Pipeline (UBP-CAP). Integrating prediction, calibration assessment, and variance decomposition modules, UBP-CAP leverages multimodal telemetry data to simultaneously predict response accuracy and quantify metacognitive bias. The work introduces the Prediction-Explanation Discrepancy Index (PEDI) to measure feature consistency between predictive and explanatory models and employs cross-validated generalized linear mixed-effects models (GLMMs) to uncover the context-dependence of calibration bias. Empirical results show that logistic regression (AUC = 0.903) outperforms LightGBM; students exhibit significantly higher calibration error (ECE = 0.109) than the model (ECE = 0.068); GLMM analysis yields an intraclass correlation coefficient (ICC) of 0.123, indicating calibration is predominantly context-driven; and PEDIcos = 0.081 reveals high alignment between prediction and explanation features.
This study addresses the limitation of the traditional Brier score, which conflates calibration and discrimination in probabilistic forecasting, thereby hindering targeted optimization. The authors propose the Manokhin probability matrix, which for the first time decouples predictive quality into two orthogonal dimensions—calibration and discrimination—by constructing a two-dimensional diagnostic framework based on the Spiegelhalter Z-statistic and the expected rank interpretation of AUC-ROC. This framework categorizes classifiers into four archetypes: Eagle, Bull, Sloth, and Mole, and reveals a theoretical asymmetry: discrimination is inherently difficult to improve, whereas calibration can be effectively post-processed. Consequently, the paper advocates a practical guideline of “optimize discrimination first, then calibrate.” Large-scale evaluation on the TabArena-v0.1 benchmark across 21 classifiers and 5 calibrators shows that Venn-Abers calibration reduces log-loss by 6.5–12.6% for Bull-type models but slightly degrades Eagle-type performance, confirming an inherent trade-off between calibration and discrimination.
This study reconciles the divergent perspectives on the Elo algorithm—viewed alternatively as a heuristic ranking system and as an online maximum likelihood estimator—and addresses the coupling between rankings and predictive models induced by estimation noise. Through theoretical analysis and data-driven insights, the authors propose a decoupling framework that introduces closed-form corrections and identification procedures for key parameters such as effective scale and home-field advantage. They further establish, for the first time, a systematic characterization of the exact and approximate relationships between binary-outcome and multi-grade Elo formulations. Leveraging stochastic gradient ascent and uniformly spaced score approximations, they develop diagnostic tools to assess convergence. Experiments on six years of FIFA men’s national team data demonstrate that the decoupled approach significantly outperforms conventional strategies that reuse rankings for prediction, while also revealing that rankings for most teams have not yet converged.