MAE-Based Self-Supervised Pretraining for Data-Efficient Medical Image Segmentation Using nnFormer

📅 2026-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of limited annotated data in medical image segmentation, where Transformer-based models like nnFormer often suffer from overfitting and unstable training due to their reliance on large labeled datasets, despite abundant unlabeled clinical images remaining underutilized. To tackle this, the study introduces a two-stage self-supervised pretraining framework by integrating Masked Autoencoders (MAE) into nnFormer: first, the encoder is pretrained on unlabeled 3D medical images using MAE to learn robust anatomical representations; then, it is fine-tuned on a small set of labeled data for segmentation. Experimental results demonstrate that this approach significantly outperforms conventional fully supervised methods, achieving notable improvements in Dice score, convergence speed, and few-shot generalization, thereby validating the efficacy and potential of self-supervised learning in medical image analysis.

Technology Category

Computer Vision: SegmentationMachine Learning: Unsupervised & Self-Supervised LearningNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Transformer architectures, including nnFormer,have demonstrated promising results in volumetric medical image segmentation by being able to capture long-range spatial interactions. Although they have high performance, these models need large quantities of labeled training data and are also likely to overfit and become training unstable. This is a serious practical problem because it is not only time-consuming but also expensive to obtain medical images that are annotated by experts. Moreover, fully supervised traditional training pipelines do not take advantage of the available large amounts of unlabeled medical imaging data that can be easily obtained in the clinics. We have solved these drawbacks by advancing the efficiency of the nnFormer with a self-supervised pretraining framework, which is based on the Masked Autoencoders (MAE). In this method, the model is pretrained on unlabeled volumetric medical images to reconstruct randomly masked parts of the input. This allows the encoder to learn meaningful anatomical and structural representations . The encoder is then further fine-tuned on a labeled dataset on the downstream segmentation task. Conducted Experiment shows that the offered method leads to a higher segmentation performance on the count of Dice score, a quicker convergence rate on the course of the fine-tuning procedure, and a superior generalization on the basis of limited labeled data . These findings validate that self-supervised learning combined with transformer-based segmentation models is an appropriate approach to the problem of data shortage in medical image analysis.
Problem

Research questions and friction points this paper is trying to address.

medical image segmentation
data efficiency
limited labeled data
unlabeled medical data
annotation scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Autoencoders
self-supervised pretraining
nnFormer
medical image segmentation
data-efficient learning
R
R. M. Krishna Sureddi
Associate Professor, Information Technology, Chaitanya Bharathi Institute of Technology, Hyderabad, India
T
T. Satyanarayana Murthy
Associate Professor, Information Technology, Chaitanya Bharathi Institute of Technology, Hyderabad, India
N
Nomula Varsha Reddy
Information Technology, Chaitanya Bharathi Institute of Technology, Hyderabad, India
A
Adi Kanishka
Information Technology, Chaitanya Bharathi Institute of Technology, Hyderabad, India
N
Nalla Manvika Reddy
Information Technology, Chaitanya Bharathi Institute of Technology, Hyderabad, India