🤖 AI Summary
Whether class imbalance correction improves the performance of clinical prediction models remains controversial. This study leverages data from the GUSTO-I clinical trial to systematically evaluate the impact of various correction strategies—including algorithm-level rebalancing, oversampling, and hybrid sampling—on model discrimination (AUC), calibration (calibration plots and MAPE), and predictive stability (Classification Instability Index, CII) across varying sample sizes. Using penalized logistic regression with 200 bootstrap replications, we find that all correction methods fail to enhance discriminative performance and instead introduce greater calibration bias, risk overestimation, and increased prediction instability. These results challenge the common practice of routinely applying class imbalance corrections in clinical modeling and, for the first time in large-scale simulations, reveal their potential harms.
📝 Abstract
Class imbalance is common when developing clinical prediction models (CPMs) and is often assumed to lead to poor predictive performance. Several methods have been proposed to correct data imbalance during CPM development. However, it remains unclear whether correcting class imbalance improves or harms CPM performance. This study investigated how imbalance correction affects classification performance and prediction stability. We simulated the development and internal validation of CPMs using penalised logistic regression under different imbalance-correction strategies, including algorithm-level rebalancing, data-level rebalancing by oversampling, and combined over- and under-sampling. The simulation dataset was derived from the GUSTO-I trial, which included 40,830 patients and 2,851 events. All imbalance-correction strategies were evaluated across sample-size scenarios ranging from 500 to 40,830. Model performance and prediction stability were assessed using 200 bootstrap resamples, including discrimination, calibration, calibration stability, mean absolute prediction error (MAPE), and classification instability index (CII). Class imbalance correction did not meaningfully improve model discrimination. Both data-level and algorithm-level correction led to miscalibration, risk overestimation, and increased prediction instability, as shown by prediction stability, MAPE, and CII plots, compared with models developed without correction. These findings suggest that class imbalance correction does not necessarily improve CPM performance and may compromise calibration and prediction stability. Class imbalance should not be treated as a pathology that automatically requires correction. In clinical prediction modelling, routine imbalance correction by default is generally not advisable.