KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

📅 2026-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Although clinical risk models often exhibit strong overall performance, they frequently demonstrate significantly disparate error rates across patient subgroups, and existing fairness audits lack reproducible stress-testing protocols. This work proposes the first systematic fairness auditing pipeline, integrating five stages: subgroup stratification, disparity measurement, mechanism diagnosis, post-hoc mitigation, and drift monitoring. The framework is rigorously evaluated on a synthetic benchmark encompassing 16 diseases, 15 social determinants of health, and predefined intersectional groups. Experiments reveal that statistical significance must be interpreted alongside minimum detectable effect sizes; calibration exhibits high variance in its impact on fairness; and mechanism diagnosis silently fails under proxy variable misspecification. The proposed group-threshold optimization consistently reduced Equal Opportunity Difference (EOD) across all 48 extrapolation scenarios, while CUSUM-based drift detection proved highly sensitive to cohort implementation, underscoring challenges in threshold transferability.
📝 Abstract
Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.
Problem

Research questions and friction points this paper is trying to address.

subgroup fairness
clinical risk models
audit pipeline
disparity measurement
reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

subgroup fairness
audit pipeline
stress testing
post-hoc mitigation
drift monitoring