🤖 AI Summary
This study addresses the lack of task-awareness and trustworthiness validation in uncertainty quantification (UQ) for AI-based glioma diagnosis. Within an MRI multi-task deep learning framework, it systematically evaluates UQ methods—including Monte Carlo Dropout, deep ensembles, and voxel-level aggregation—for estimating predictive, aleatoric, and epistemic uncertainties. The work proposes a task-aware UQ evaluation strategy, investigates the uncertainty interaction mechanisms between segmentation and classification tasks, and constructs composite trust scores. Results demonstrate that moderate dropout rates achieve optimal calibration, offering practical guidelines for clinical deployment. However, the composite trust scores do not exhibit significant advantages over single-task classification uncertainty.
📝 Abstract
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.