UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

šŸ“… 2026-07-31
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work addresses the challenges of ultrasound image segmentation—namely speckle noise, low contrast, and ambiguous boundaries—and overcomes the limited generalizability of existing methods that rely on task-specific designs or expert-provided prompts. We propose UltraSAM3, the first text-driven, concept-oriented foundation model for universal ultrasound segmentation, built upon the SAM3 architecture and adapted to the ultrasound modality. Trained on a large-scale dataset comprising 37 sources and 13 anatomical categories, UltraSAM3 leverages image–mask–concept triplets to achieve cross-organ and cross-lesion semantic alignment. By integrating a natural language instruction interpreter, the model significantly enhances clinical interactivity and robustness under complex textual commands. Experiments demonstrate that UltraSAM3 consistently outperforms current text- or concept-driven approaches across multi-organ benchmarks, external test sets, and visual enhancement scenarios.
šŸ“ Abstract
Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound image segmentation important. However, ultrasound images differ substantially from CT, MRI, and other medical imaging modalities, as they are often affected by speckle noise, low contrast, acoustic shadows and ambiguous boundaries. Existing ultrasound segmentation methods are still mainly limited to task-specific models or visual-prompt-based foundation models, which are either tailored to particular tasks or require expert-provided visual prompts, making them inconvenient for flexible clinical use. To address these challenges, we propose UltraSAM3, a concept-driven foundation model for universal ultrasound image segmentation. Unlike conventional models, UltraSAM3 enables text-based target specification by adapting SAM3 to ultrasound-specific image--mask--concept triplets. The model is trained on a large-scale ultrasound segmentation corpus covering 37 public datasets and 13 anatomical categories, allowing it to align ultrasound visual patterns with clinically meaningful concepts across diverse organs and lesions. To further improve usability under realistic clinical interaction, we propose an instruction-guided agent that parses complex natural language queries into concise ultrasound concept prompts for UltraSAM3. Extensive experiments demonstrate that UltraSAM3 consistently outperforms representative concept- and text-driven biomedical segmentation models on multi-organ ultrasound benchmarks, external datasets, and visual-prompt-enhanced settings. Moreover, the agent improves segmentation robustness for complex user instructions. These results indicate that ultrasound-specific concept adaptation is effective for building generalizable and interactive ultrasound segmentation foundation models.
Problem

Research questions and friction points this paper is trying to address.

ultrasound image segmentation
foundation model
concept-driven
clinical usability
speckle noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

concept-driven
ultrasound segmentation
foundation model
text-to-segmentation
instruction-guided agent
šŸ”Ž Similar Papers
No similar papers found.