🤖 AI Summary
Existing calibration assessment methods for multicenter clinical prediction models overlook inter-center heterogeneity in calibration, leading to potentially misleading aggregate evaluations.
Method: We propose three novel calibration plot methods—CG-C, 2MA-C, and MIX-C—that explicitly incorporate clustering structure by integrating random-effects modeling with spline-based smoothing. These methods jointly estimate population-average and center-specific calibration curves.
Contribution/Results: The framework enables quantification of calibration heterogeneity, robust inference for small-sample centers, and simultaneous estimation of confidence and prediction intervals. In a multicenter validation study for ovarian tumor malignancy risk prediction, MIX-C most closely approximated the true center-specific calibration curves, while 2MA-C (with splines) achieved optimal prediction interval coverage. All methods are open-source, modular, and readily deployable. Collectively, they establish a unified, robust, and interpretable statistical framework for calibration assessment in multicenter prediction modeling.
📝 Abstract
Evaluation of clinical prediction models across multiple clusters, whether centers or datasets, is becoming increasingly common. A comprehensive evaluation includes an assessment of the agreement between the estimated risks and the observed outcomes, also known as calibration. Calibration is of utmost importance for clinical decision making with prediction models and it may vary between clusters. We present three approaches to take clustering into account when evaluating calibration. (1) Clustered group calibration (CG-C), (2) two stage meta-analysis calibration (2MA-C) and (3) mixed model calibration (MIX-C) can obtain flexible calibration plots with random effects modelling and providing confidence and prediction intervals. As a case example, we externally validate a model to estimate the risk that an ovarian tumor is malignant in multiple centers (N = 2489). We also conduct a simulation study and synthetic data study generated from a true clustered dataset to evaluate the methods. In the simulation and the synthetic data analysis MIX-C gave estimated curves closest to the true overall and center specific curves. Prediction interval was best for 2MA-C with splines. Standard flexible calibration worked likewise in terms of calibration error when sample size is limited. We recommend using 2MA-C (splines) to estimate the curve with the average effect and the 95% PI and MIX-C for the cluster specific curves, specially when sample size per cluster is limited. We provide ready-to-use code to construct summary flexible calibration curves with confidence and prediction intervals to assess heterogeneity in calibration across datasets or centers.