🤖 AI Summary
This study addresses the bias in Cox proportional hazards model estimates arising from measurement error in AI-extracted covariates, a setting where downstream users only have access to the extracted data and limited calibration summary statistics. Within a multivariate calibration framework, the work provides the first decomposition of Cox model bias into a dominant, calibratable component and higher-order residual terms. Building on this insight, the authors propose a post-processing correction method that relies solely on calibration summary statistics and can be directly applied to outputs from standard Cox regression software. The approach is accompanied by uncertainty-adjusted confidence intervals and sensitivity diagnostic tools. Empirical evaluations on synthetic data demonstrate substantial bias reduction, with near-nominal coverage maintained even under mild violations of the linear calibration assumption. The paper also recommends a minimal set of calibration statistics that data providers should report to enable effective bias correction.
📝 Abstract
Large-scale observational studies increasingly rely on AI pipelines to extract structured variables from unstructured clinical records. A common workflow separates the data vendor, who validates extraction accuracy with a gold-standard sample, from the downstream researcher, who receives only the extracted dataset and summary accuracy statistics. We develop a bias-correction framework for the Cox proportional hazards model when covariates are subject to AI extraction error. Within a unified multivariate calibration framework, we show that the naive Cox estimator's bias decomposes into a leading-order calibration term and a second-order residual that vanishes as extraction accuracy improves. The leading-order term yields a corrected estimator that operates as a post-hoc matrix multiplication on the output of any standard Cox software. We further derive bias-adjusted confidence intervals that incorporate calibration uncertainty and a sensitivity diagnostic for assessing whether the neglected residual could materially affect inference. Synthetic data experiments with cross-dependent extraction errors and controlled nonlinear calibration violations confirm that the correction substantially reduces bias and achieves near-nominal coverage even under mild violations of the linear calibration assumption. The framework yields a concrete reporting specification: a short list of summary statistics that data vendors should provide alongside any AI-extracted covariate dataset used in survival analysis.