🤖 AI Summary
This study addresses the confidence calibration bias caused by missing modalities in multimodal brain tumor segmentation by proposing the MMA-LTS method. Challenging the conventional assumption that task difficulty depends solely on the number of missing modalities, this work reveals that predictive uncertainty is instead determined by the specific combination of absent modalities. Accordingly, MMA-LTS introduces learnable modality-availability tokens, voxel-wise difficulty scores, and a local temperature scaling mechanism to achieve spatially adaptive, voxel-level post-hoc calibration. Experiments on the BraTS and FeTS datasets demonstrate that MMA-LTS significantly improves calibration accuracy while maintaining state-of-the-art segmentation performance, thereby effectively enhancing the clinical trustworthiness of the model.
📝 Abstract
Multimodal brain tumor segmentation typically leverages multiple MRI modalities, yet incomplete modality acquisition is common in clinical practice due to protocol heterogeneity and scan failures. Although recent methods maintain segmentation accuracy under missing modality conditions, they frequently overlook prediction reliability, leading to miscalibrated confidence estimates that hinder clinical adoption. Existing calibration techniques are largely modality-agnostic or assume that prediction difficulty decreases monotonically as additional modalities become available. However, in brain tumor segmentation, prediction difficulty depends primarily on which modalities are absent rather than how many, leading to combination-specific and spatially heterogeneous calibration errors. To address this, we propose Missing Modality-Aware Local Temperature Scaling (MMA-LTS), a post-hoc voxel-wise confidence calibration method. It estimates a spatially adaptive temperature field conditioned on a modality-availability learnable token and a voxel-wise difficulty score. Experiments on BraTS 2020 and FeTS 2024 show that MMA-LTS improves calibration while preserving the segmentation accuracy of state-of-the-art models across diverse missing-modality scenarios, thereby enhancing trustworthiness toward clinical deployment.