🤖 AI Summary
This study addresses the unexplored robustness of promptable concept segmentation models, such as SAM3, and the limited cross-prompt transferability of existing adversarial attacks. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack framework. AdvPCS introduces a minimax prompt optimization strategy to enhance the diversity and confidence of point, box, and text prompts. Furthermore, it employs a bilevel optimization framework to jointly execute global-local perceptual deception and temporal memory misalignment attacks. Experimental results demonstrate that the generated universal perturbations exhibit strong cross-frame generalization capabilities. On the SA-CO dataset, AdvPCS reduces the average mIoU of various target models to below 5%, effectively overcoming the limitations associated with single-prompt adversarial attacks.
📝 Abstract
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global-local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5% under text prompts, demonstrating strong attack ability.