Examining marginal properness in the external validation of survival models with squared and logarithmic losses

📅 2022-12-10
📈 Citations: 3
Influential: 0
📄 PDF

career value

150K/year
🤖 AI Summary
This paper addresses the theoretical validity of two widely used external validation metrics for survival analysis models—Integrated Survival Brier Score (ISBS) and Right-Censored Log-Likelihood (RCLL)—by introducing “marginal propriety” as a novel formal criterion for scoring rule appropriateness. We prove theoretically that neither metric satisfies marginal propriety. However, Monte Carlo simulations and extensive experiments across diverse right-censored survival modeling scenarios demonstrate that RCLL consistently satisfies this property empirically, while ISBS exhibits only negligible violations under extremely small sample sizes, remaining robust in practice. This reveals a critical dissociation between theoretical impropriety and empirical robustness—a key insight with important implications for metric selection and design. Building on this finding, we propose a new class of loss-function frameworks for survival prediction, grounded in marginal propriety, thereby providing both theoretical guidance and practical foundations for developing future survival scoring rules.
📝 Abstract
Scoring rules promote rational and honest decision-making, which is important for model evaluation and becoming increasingly important for automated procedures such as `AutoML'. In this paper we survey common squared and logarithmic scoring rules for survival analysis, with a focus on their theoretical and empirical properness. We introduce a marginal definition of properness and show that both the Integrated Survival Brier Score (ISBS) and the Right-Censored Log-Likelihood (RCLL) are theoretically improper under this definition. We also investigate a new class of losses that may inform future survival scoring rules. Simulation experiments reveal that both the ISBS and RCLL behave as proper scoring rules in practice. The RCLL showed no violations across all settings, while ISBS exhibited only minor, negligible violations at extremely small sample sizes, suggesting one can trust results from historical experiments. As such we advocate for both the RCLL and ISBS in external validation of models, including in automated procedures. However, we note practical challenges in estimating these losses including estimation of censoring distributions and densities; as such further research is required to advance development of robust and honest evaluation in survival analysis.
Problem

Research questions and friction points this paper is trying to address.

Evaluating theoretical and empirical properness of survival scoring rules
Investigating marginal properness in ISBS and RCLL survival metrics
Addressing practical challenges in robust survival model validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introducing marginal properness definition for survival models
Evaluating ISBS and RCLL as practically proper scoring rules
Proposing new loss class for future survival scoring rules
🔎 Similar Papers
No similar papers found.
R
R. Sonabend
University of Kaiserslautern-Landau, Germany; Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI), Germany
J
John Zobolas
Oslo University Hospital, Norway
Philipp Kopper
Philipp Kopper
University of Munich
Machine LearningDeep LearningStatisticsComputational Statistics
L
Lukas Burk
LMU Munich, Germany; Munich Center for Machine Learning (MCML), Germany; Leibniz Institute for Prevention Research and Epidemiology - BIPS, Germany; University of Bremen, Germany
A
Andreas Bender
LMU Munich, Germany; Munich Center for Machine Learning (MCML), Germany