π€ AI Summary
Classic Brier score neglects clinical decision impact, limiting its ability to assess the real-world utility of risk prediction models. To address this, we propose a clinical-utility-oriented weighted Brier score framework that integrates decision-sensitive weights to jointly quantify predictive accuracy and costβbenefit trade-offs. Our work is the first to decompose this score into discriminative and calibration components and to establish a theoretical connection with the H-measure, thereby bridging the gap between conventional calibration assessment and decision-theoretic evaluation. Leveraging decision modeling, weighted generalization, decomposition analysis, and rigorous theoretical derivation, we validate the framework on the Prostate Active Surveillance Study (PASS) cohort. Results demonstrate that the proposed score exhibits high sensitivity to clinically relevant risk thresholds and significantly outperforms both the classic Brier score and AUC. It thus serves as a comprehensive pre-deployment evaluation metric for clinical risk models.
π Abstract
As advancements in novel biomarker-based algorithms and models accelerate disease risk prediction and stratification in medicine, it is crucial to evaluate these models within the context of their intended clinical application. Prediction models output the absolute risk of disease; subsequently, patient counseling and shared decision-making are based on the estimated individual risk and cost-benefit assessment. The overall impact of the application is often referred to as clinical utility, which received significant attention in terms of model assessment lately. The classic Brier score is a popular measure of prediction accuracy; however, it is insufficient for effectively assessing clinical utility. To address this limitation, we propose a class of weighted Brier scores that aligns with the decision-theoretic framework of clinical utility. Additionally, we decompose the weighted Brier score into discrimination and calibration components, examining how weighting influences the overall score and its individual components. Through this decomposition, we link the weighted Brier score to the $H$ measure, which has been proposed as a coherent alternative to the area under the receiver operating characteristic curve. This theoretical link to the $H$ measure further supports our weighting method and underscores the essential elements of discrimination and calibration in risk prediction evaluation. The practical use of the weighted Brier score as an overall summary is demonstrated using data from the Prostate Cancer Active Surveillance Study (PASS).