ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model

📅 2026-07-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of low lesion-to-background contrast in non-contrast CT (NCCT), which limits segmentation accuracy for proton therapy planning. The authors propose the first approach to integrate cross-modal visual prompts into the Segment Anything Model (SAM), leveraging contrast-enhanced MRI as guidance. By incorporating a visual prompt encoder and a visual-guidance cross-attention module, the method effectively fuses multimodal features and fine-tunes the mask decoder in a parameter-efficient manner. This strategy substantially enhances the representation of low-contrast lesions in NCCT, achieving superior performance over both U-Net and the original SAM on liver lesion segmentation, with improved accuracy and robustness.
📝 Abstract
In proton therapy planning, respiratory-gated non-contrast CT (NCCT) is commonly used for lesion segmentation; however, accurate delineation remains challenging due to low lesion-to-background contrast. Although learning-based methods have shown strong performance, they often struggle with non-contrast image segmentation. Inspired by clinical practice, where contrast-enhanced MRI is referenced to delineate lesions on NCCT, we propose ViPSAM, a visual prompting framework that leverages complementary cross-modality information. Built upon the Segment Anything Model (SAM), ViPSAM introduces a visual prompt encoder to extract guidance features from contrast-enhanced images and a visual-guided cross-attention module to integrate non-contrast and contrast-enhanced features, thereby enhancing lesion-relevant representations in low-contrast regions. The mask decoder is further adapted in a parameter-efficient manner to utilize visual prompts effectively. We evaluate the proposed method on liver lesion segmentation using NCCT acquired for proton therapy. Experimental results demonstrate that ViPSAM outperforms representative U-Net- and SAM-based methods, indicating that cross-modality visual prompting enables more robust and accurate segmentation in non-contrast images.
Problem

Research questions and friction points this paper is trying to address.

medical image segmentation
non-contrast CT
low lesion-to-background contrast
cross-modality
proton therapy planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual prompting
cross-modality
Segment Anything Model
medical image segmentation
parameter-efficient adaptation
🔎 Similar Papers
No similar papers found.
S
San Lee
Department of Artificial Intelligence, Sungkyunkwan University, Republic of Korea
N
Nalee Kim
Department of Radiation Oncology, Samsung Medical Center, Sungkyunkwan University School of Medicine, Republic of Korea
J
Jeong Il Yu
Department of Radiation Oncology, Samsung Medical Center, Sungkyunkwan University School of Medicine, Republic of Korea
H
Hee Chul Park
Department of Radiation Oncology, Samsung Medical Center, Sungkyunkwan University School of Medicine, Republic of Korea
Boah Kim
Boah Kim
Sungkyunkwan University (SKKU)
Artificial intelligenceMedical imagingComputer visionLarge language model