Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unexplored robustness of promptable concept segmentation models, such as SAM3, and the limited cross-prompt transferability of existing adversarial attacks. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack framework. AdvPCS introduces a minimax prompt optimization strategy to enhance the diversity and confidence of point, box, and text prompts. Furthermore, it employs a bilevel optimization framework to jointly execute global-local perceptual deception and temporal memory misalignment attacks. Experimental results demonstrate that the generated universal perturbations exhibit strong cross-frame generalization capabilities. On the SA-CO dataset, AdvPCS reduces the average mIoU of various target models to below 5%, effectively overcoming the limitations associated with single-prompt adversarial attacks.
📝 Abstract
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global-local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5% under text prompts, demonstrating strong attack ability.
Problem

Research questions and friction points this paper is trying to address.

Adversarial Attack
Promptable Concept Segmentation
Cross-Prompt Transferability
Robustness
SAM3
Innovation

Methods, ideas, or system contributions that make the work stand out.

Universal Adversarial Attack
Cross-Prompt Transferability
Promptable Concept Segmentation
Min-Max Optimization
Temporal Memory Misalignment
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Ziqi Zhou
Ziqi Zhou
Huazhong University of Science and Technology (HUST)
Trustworthy AI
Y
Yifan Hu
School of Cyber Science and Engineering, Huazhong University of Science and Technology
Y
Yufei Song
School of Cyber Science and Engineering, Huazhong University of Science and Technology
H
Haowen Jiang
School of Computer Science and Technology, Huazhong University of Science and Technology
Xianlong Wang
Xianlong Wang
Ph.D. student, City University of Hong Kong
Trustworthy LLM/VLMEmbodied AIUnlearnable Example3D Point CloudPoisoning/Adversarial Attack
Shengshan Hu
Shengshan Hu
School of CSE, Huazhong University of Science and Technology (HUST)
AI SecurityEmbodied AIAutonomous Driving
Dezhong Yao
Dezhong Yao
Huazhong University of Science and Technology
Distributed Machine LearningEdge ComputingFederated Learning
L
Leo Yu Zhang
School of Information and Communication Technology, Griffith University