🤖 AI Summary
This study addresses the unclear robustness of existing 3D medical foundation models under MRI artifacts, a critical limitation for clinical reliability. Within a unified framework, it systematically evaluates the representation stability of five prominent self-supervised 3D encoders—including 3DINO and BrainIAC—across diverse image-domain and frequency-domain artifacts. The analysis integrates linear CKA, RankMe, UMAP visualization, and downstream segmentation consistency metrics. Findings reveal that 3DINO exhibits the most robust representations; notably, large-scale or domain-specific pretraining does not inherently confer artifact invariance. While artifacts frequently distort the geometric structure of learned representations, they do not necessarily induce dimensional collapse. This work establishes a new benchmark and offers empirical insights for assessing the robustness of medical foundation models.
📝 Abstract
Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood. We present a controlled evaluation of representation robustness across five pretrained 3D encoders spanning different architectures, objectives, pretraining domains, and dataset scales. Using BraTS-Africa cases with four MRI sequences, we generate seven frequency- and image-domain artifacts at five predefined corruption settings. Robustness is assessed using linear centered kernel alignment (CKA), RankMe, and UMAP, complemented by an independent segmentation-consistency analysis. We find that robustness is strongly model- and artifact-dependent. 3DINO exhibits the most consistently stable representations, while BrainIAC is highly sensitive to several corruptions; NeuroVFM, BrainFM, and Neuro-SimCLR show intermediate but distinct artifact-specific profiles. Across many conditions, CKA decreases substantially while RankMe remains comparatively stable, indicating that artifacts often distort representation geometry without causing dimensional collapse. Segmentation consistency also degrades under corruption, particularly for ghosting and Rician noise, but aligns only partially with representation-level robustness. These findings show that larger-scale or domain-specific pretraining alone does not guarantee artifact invariance and motivate explicit robustness evaluation before deploying 3D foundation models in heterogeneous MRI settings.