Score
Designs and evaluates self-supervised (SSL) pretraining pipelines for MRI data that learn general-purpose feature encoders from unlabeled scans, including choices of architectures, pretext tasks or contrastive objectives, augmentations, and transfer procedures. Builds and analyzes how these pretrained encoders reduce labeled-data requirements and support downstream MRI tasks such as classification, segmentation, or regression by measuring representation quality and label efficiency.
This study addresses the performance bottleneck in self-supervised learning (SSL) for medical imaging, which arises from misalignment between pretraining objectives and downstream clinical tasks. Synthesizing 75 studies from 2017 to 2025, it introduces a novel task-alignment perspective to categorize SSL methods into four paradigms: contrastive, non-contrastive predictive, generative reconstruction, and hybrid approaches. Guided by the PRISMA framework and grounded in modality-specific characteristics of medical images, the work conducts performance attribution analysis, proposes principles for co-designing modalities and pretext tasks, and distills practical guidelines for common downstream tasks such as classification and segmentation. Findings indicate substantial gains from SSL under low-label and few-shot settings; contrastive methods excel in classification, while generative and spatial prediction strategies benefit segmentation, with hybrid approaches offering the most balanced performance. The paper concludes by advocating standardized evaluation protocols and pathology-aware pretraining as key future directions.
To address three key bottlenecks in self-supervised learning (SSL) for 3D medical image segmentation—limited pretraining data scale, architectural mismatch with 3D convolutional networks, and insufficient evaluation—this paper introduces the first end-to-end SSL pretraining framework specifically designed for large-scale 3D brain MRI. Leveraging 39,000 multi-center brain MRI scans, we adapt the Masked Autoencoder (MAE) to 3D CNNs via a residual encoder U-Net architecture and deeply integrate nnU-Net’s preprocessing, augmentation, and inference pipelines. We conduct systematic cross-center evaluation across five development and eight test datasets, demonstrating substantially improved generalization robustness. Our method achieves an average Dice score gain of ~3.0% over state-of-the-art SSL approaches and the strong supervised baseline nnU-Net, establishing new SOTA in 3D medical image segmentation. The code and pretrained models are publicly available.
This work proposes a novel self-supervised pretraining paradigm for 3D MRI that moves beyond treating scans as static slices or voxel sets, which neglects their inherent spatial structure and navigational dynamics. Instead, the method transforms 3D MRI volumes into controllable sequences of 2D slice renderings, constructing video–action pairs that encode positional, orientational, and scale-related actions to serve as self-supervisory signals. By integrating a slice observation encoder with an action-conditioned latent dynamics model, 3D anatomical understanding is framed as a dynamic sequence prediction task. Experiments demonstrate that this approach significantly outperforms static reconstruction baselines, encoder-only pretraining schemes, and dynamics-based variants without action alignment across multiple anatomical and spatial downstream tasks, thereby validating its effectiveness and conceptual novelty.
Current 3D medical self-supervised learning (SSL) suffers from inconsistent dataset scales and diversity, heterogeneous model architectures, and a lack of standardized downstream evaluation protocols, hindering fair methodological comparison. To address this, we introduce the first large-scale, publicly available 3D brain MRI pretraining dataset (114K scans) and establish a unified benchmark—standardizing data, architectures (CNN/ViT), and downstream tasks (multi-center segmentation)—to systematically evaluate mainstream SSL paradigms (contrastive learning and masked modeling). Our contributions are threefold: (1) the largest open-source 3D brain MRI pretraining dataset to date; (2) the first standardized 3D medical SSL benchmark; and (3) fully open-sourced training framework, pretrained models, and reproducible code. Experiments demonstrate that SSL pretraining significantly outperforms end-to-end trained nnU-Net ResEnc-L baselines and substantially improves segmentation performance in low-data regimes, establishing a new state-of-the-art practice for 3D medical SSL.
This work proposes SSPFormer, a self-supervised representation learning method for medical imaging that addresses the challenges of adapting pretrained Transformers to the anatomical specificity of MRI, as well as the scarcity and privacy constraints of medical data. SSPFormer innovatively integrates inverse frequency-domain projection masking—which prioritizes reconstruction of high-frequency anatomical regions—with physiologically plausible frequency-weighted FFT noise augmentation. This enables structure-aware and artifact-robust feature learning directly from unlabeled raw MRI data. Built upon a Transformer architecture, SSPFormer achieves state-of-the-art performance across multiple tasks, including segmentation, super-resolution, and denoising, significantly enhancing MRI detail fidelity and demonstrating strong clinical applicability.
This study systematically evaluates the transferability of nine self-supervised learning (SSL) methods for 3D medical image segmentation, with a particular focus on label-scarce scenarios. Leveraging a unified CT pretraining protocol and the SwinUNETR architecture, the authors fine-tune models across nine diverse CT and MRI segmentation tasks, analyzing convergence speed, cross-modality transferability, and feature reuse patterns through Dice scores and Centered Kernel Alignment (CKA). This large-scale, first-of-its-kind comparative study reveals that SMIT achieves superior overall performance in terms of accuracy, convergence speed, and few-shot stability. Furthermore, mask image modeling (MIM) and self-distillation consistently outperform contrastive learning and rotation prediction, underscoring the critical advantage of local representation learning under data-limited conditions.
This work addresses the lack of systematic investigation into self-supervised foundation models for MRI-based disease detection, a domain where prior efforts have predominantly focused on dense prediction tasks such as segmentation. The study systematically compares two prominent paradigms—Masked Autoencoders (MAE) and Joint-Embedding Predictive Architectures (JEPA)—proposing a spectral-domain reconstruction loss to enhance MAE’s sensitivity to fine anatomical structures and introducing variance–covariance regularization in JEPA to encourage decorrelated latent representations. For the first time, it reveals the critical alignment between self-supervised pretraining objectives and the structural characteristics of downstream discriminative signals, enabling contrast-agnostic pretraining on single-contrast, heterogeneous 3D brain MRI without modality concatenation. Experiments across five disease detection tasks demonstrate that spectrally supervised MAE achieves superior performance, underscoring the importance of aligning pretraining objectives with downstream task structure.
This work addresses the challenges of optimization instability and artifact generation in zero-shot self-supervised methods for single-coil undersampled MRI reconstruction, which often arise due to limited supervisory signals. To enhance both stability and reconstruction accuracy, the authors propose a physics-driven zero-shot framework that synergistically integrates three mechanisms: coil sensitivity–guided dynamic image priors, k-space self-consistency regularization based on SPIRiT kernel modeling, and non-local self-similarity exploitation. Notably, the method requires no additional training data and achieves state-of-the-art performance on the FastMRI dataset, substantially narrowing the performance gap between zero-shot and supervised learning approaches—particularly under high acceleration factors.
Clinical brain MRI analysis is hindered by data heterogeneity, high noise levels, and the prohibitive cost of annotations, which impede the clinical deployment of automated models. To address this, this work introduces the FOMO25 challenge, leveraging the large-scale unlabeled clinical dataset FOMO60K to systematically evaluate the generalization capabilities of self-supervised foundation models under few-shot and out-of-distribution settings across three tasks: infarction classification, meningioma segmentation, and brain age regression. Innovatively benchmarking model performance on real-world clinical workflow data, the study reveals task-dependent effects of different self-supervision objectives and demonstrates that even modestly sized pre-trained models can achieve strong performance. Results show that self-supervised pre-training substantially enhances generalization, with the best out-of-distribution model surpassing in-domain supervised baselines, while increased model scale and training duration yield no consistent gains.
This work addresses the challenge of limited annotated data in medical image segmentation, where Transformer-based models like nnFormer often suffer from overfitting and unstable training due to their reliance on large labeled datasets, despite abundant unlabeled clinical images remaining underutilized. To tackle this, the study introduces a two-stage self-supervised pretraining framework by integrating Masked Autoencoders (MAE) into nnFormer: first, the encoder is pretrained on unlabeled 3D medical images using MAE to learn robust anatomical representations; then, it is fine-tuned on a small set of labeled data for segmentation. Experimental results demonstrate that this approach significantly outperforms conventional fully supervised methods, achieving notable improvements in Dice score, convergence speed, and few-shot generalization, thereby validating the efficacy and potential of self-supervised learning in medical image analysis.