🤖 AI Summary
This study challenges the implicit “more modalities, better performance” assumption in multimodal deep learning (MDL) for computational pathology, focusing on survival time prediction for prostate cancer biochemical recurrence. We address the problem that integrating low-performing modalities—such as histopathology images, MRI, and clinical variables—can introduce noise and degrade predictive accuracy. To mitigate this, we propose a performance-guided multimodal fusion strategy: only modalities demonstrating strong independent prognostic value in survival analysis are selected for fusion. Experimental results show that selective fusion significantly improves both the concordance index (C-index) and Brier score, whereas inclusion of low-performing modalities consistently harms performance. To our knowledge, this is the first systematic investigation validating the critical impact of modality quality on MDL efficacy in survival prediction. Our work establishes modality selection—not merely fusion—as a fundamental step for enhancing robustness and reliability in multimodal survival modeling.
📝 Abstract
Multimodal deep learning (MDL) has emerged as a transformative approach in computational pathology. By integrating complementary information from multiple data sources, MDL models have demonstrated superior predictive performance across diverse clinical tasks compared to unimodal models. However, the assumption that combining modalities inherently improves performance remains largely unexamined. We hypothesise that multimodal gains depend critically on the predictive quality of individual modalities, and that integrating weak modalities may introduce noise rather than complementary information. We test this hypothesis on a prostate cancer dataset with histopathology, radiology, and clinical data to predict time-to-biochemical recurrence. Our results confirm that combining high-performing modalities yield superior performance compared to unimodal approaches. However, integrating a poor-performing modality with other higher-performing modalities degrades predictive accuracy. These findings demonstrate that multimodal benefit requires selective, performance-guided integration rather than indiscriminate modality combination, with implications for MDL design across computational pathology and medical imaging.