🤖 AI Summary
This study addresses the insufficient uncertainty calibration and inter-class bias in 3D medical segmentation models under distribution shifts by proposing the CA-WCP method. This approach integrates latent-space density ratio weighting with conformal prediction, introducing class-specific asymmetric factors and directional quantiles to overcome traditional symmetry assumptions and enable class-aware volumetric risk quantification. Furthermore, it leverages multimodal large language model prompt engineering to generate uncertainty-conditioned radiology reports. Experimental results demonstrate that the proposed method effectively guarantees coverage rates across all classes while reducing prediction interval widths by 8–14%, significantly enhancing both the clinical reliability and interpretability of the models.
📝 Abstract
Reliable volumetric segmentation is critical for clinical diagnostics, yet foundation models such as MedSAM remain deterministic and lack calibrated uncertainty under distribution shift. Existing conformal prediction methods offer statistical guarantees but are frequently applied in 2D and assume symmetric error distributions, so they do not capture the class-specific biases that arise in 3D multi-class segmentation. We propose Class-Aware Asymmetric Weighted Conformal Prediction (CA-WCP), which combines latent-space density-ratio weighting for covariate shift with directional quantiles for the lower and upper volume bounds, and scales each bound by a class-specific asymmetry factor derived from validation-set false-positive and false-negative rates. We prove that CA-WCP retains the weighted-exchangeability marginal coverage guarantee for every class, and we evaluate it on 3D brain tumor segmentation (BraTS 2020) and on a synthetic multi-organ CT benchmark constructed under covariate shift. On both benchmarks the 95\% Clopper--Pearson interval for the observed coverage of CA-WCP contains the nominal 90\% level for every semantic class, while interval width is reduced by 8--14\% relative to symmetric weighted conformal prediction. We further encode the calibrated intervals into structured prompts for a multimodal large language model to produce uncertainty-conditioned radiology reports, linking distribution-shift-aware uncertainty quantification to interpretable clinical communication.