🤖 AI Summary
This study addresses the heavy reliance on expert annotations in MRI-based grading of lumbar degenerative diseases by proposing a label-efficient learning paradigm. The approach first leverages automated segmentation tools to generate pseudo-labels for vertebrae, intervertebral discs, and spinal canals, followed by pretraining a 3D ResNet encoder on approximately 2,000 multi-center 3D MRI scans. A lightweight task-specific head is then fine-tuned using varying proportions of real clinical grades. Experimental results demonstrate that the method achieves performance close to fully supervised baselines with only 20% of the true labels, substantially reducing annotation burden—particularly benefiting the assessment of low-prevalence and spatially localized pathologies. The segmentation pretraining attains a Dice score of 0.94 and consistently improves average ROC-AUC across all labeling ratios, providing the first empirical validation of segmentation-based pretraining’s efficacy for degenerative grading tasks.
📝 Abstract
Automated assessment of degenerative pathology in the lumbar spine on magnetic resonance imaging (MRI) requires access to large-scale datasets of expert-annotated radiological gradings. In contrast, segmentation pseudo-labels can be generated by automated tools at negligible radiologist cost. We examine whether pre-training on segmentation can effectively replace a fraction of the manual grading annotations required for downstream supervision. We pre-train a 3D ResNet encoder to segment the vertebrae, intervertebral discs (IVDs), and the spinal canal, then fine-tune lightweight task-specific grading heads using different proportions of the available training data, ranging from $10\%$ to $100\%$. On a multicentre dataset of ${\sim}2{,}000$ subjects across 11 pathologies, segmentation pre-training, achieving a Dice score of $0.94$ against pseudo-labels, improved the task-averaged (macro) one-vs-rest ROC-AUC at all proportions. With only 20\% of grading labels after pre-training, the method achieved near full-supervision performance, with the largest gains observed for either low-prevalence or spatially grounded pathologies.