🤖 AI Summary
This study addresses poor generalizability in MRI-based femur segmentation—caused by small anatomical structures and dataset bias—by systematically evaluating U-Net, Attention U-Net, U-KAN (a Kolmogorov–Arnold network-based architecture), and the prompt-driven vision transformer SAM 2 across 11,164 clinical MRI scans. It represents the first comparative assessment of CNNs versus prompt-guided transformer architectures for this task, with segmentation accuracy quantified via Dice scores (0.932–0.954). Results indicate that Attention U-Net achieves the best overall performance, while U-KAN demonstrates significantly improved accuracy in fine-grained regions—particularly the intercondylar ridge—highlighting superior modeling capability for small structures. The findings provide empirically grounded guidance for model selection and architectural design in medical image segmentation of subtle anatomical features.
📝 Abstract
Convolutional neural networks like U-Net excel in medical image segmentation, while attention mechanisms and KAN enhance feature extraction. Meta's SAM 2 uses Vision Transformers for prompt-based segmentation without fine-tuning. However, biases in these models impact generalization with limited data. In this study, we systematically evaluate and compare the performance of three CNN-based models, i.e., U-Net, Attention U-Net, and U-KAN, and one transformer-based model, i.e., SAM 2 for segmenting femur bone structures in MRI scan. The dataset comprises 11,164 MRI scans with detailed annotations of femoral regions. Performance is assessed using the Dice Similarity Coefficient, which ranges from 0.932 to 0.954. Attention U-Net achieves the highest overall scores, while U-KAN demonstrated superior performance in anatomical regions with a smaller region of interest, leveraging its enhanced learning capacity to improve segmentation accuracy.