🤖 AI Summary
This work addresses the challenges of limited data availability, substantial modality discrepancies, and insufficient spatial context in point-based prompts that hinder accuracy in medical image segmentation. Building upon MedSAM, the authors introduce a lightweight Box Predictor module that generates an approximate bounding box from a single click to enhance spatial guidance, alongside a two-stage fine-tuning strategy to improve cross-modal generalization. The proposed module adds only 1.6 million parameters with negligible inference overhead, yet significantly enriches the spatial semantics of point prompts. Evaluated on four diverse datasets—FLARE22, BRISC, BUSI, and LungSegDB—the method achieves Dice scores of 0.93, 0.88, 0.89, and 0.98, respectively, demonstrating marked improvements in both segmentation accuracy and robustness.
📝 Abstract
Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities. While foundation models like the Segment Anything Model (SAM) show promise, they often struggle with medical images without specific adaptation. Moreover, point prompts, despite being the most natural form of user interaction, provide insufficient spatial context for reliable segmentation, particularly when target structures are irregular or poorly contrasted. In this paper, we propose an enhanced segmentation framework that integrates a lightweight Box Predictor module into the MedSAM architecture. The Box Predictor estimates an approximate bounding box from a single user click using localized image embedding features, providing spatial guidance that reduces the ambiguity of point prompts, while introducing only 1.6M additional parameters and negligible inference overhead. We introduce a two-stage training pipeline where the Box Predictor is trained independently before being integrated into MedSAM. To validate the generalization capability of our method, we conduct extensive evaluations on four diverse datasets (FLARE22, BRISC, BUSI, LungSegDB) spanning distinct imaging modalities, including CT, MRI, and Ultrasound. Our method improves segmentation accuracy and robustness across varied anatomical structures and imaging domains, achieving Dice scores of 0.89 (BUSI), 0.93 (FLARE22), 0.88 (BRISC), and 0.98 (LungSegDB). Code is available at https://github.com/Amirhosseinmovahedi/MedSAM-BoxPredictor