One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of insufficient attack transferability in the black-box safety evaluation of frontier multimodal large models by proposing the O-Attack framework. The method identifies and anchors underutilized high-level cross-modal aligned semantic representations within surrogate models, optimizes perturbations via a semantic consensus mechanism, and generates a single adversarial image through a semantically conditioned progressive expansion strategy to achieve efficient black-box transfer attacks. Experimental results demonstrate that this framework elevates the attack success rate against 24 mainstream models, including GPT-5.4, to approximately 80%, significantly outperforming existing state-of-the-art methods while maintaining both high transferability and strong stealthiness.
📝 Abstract
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.
Problem

Research questions and friction points this paper is trying to address.

Adversarial attacks
Multimodal large language models
Black-box transferability
Robustness evaluation
Safety risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box adversarial attack
Transferability
Multimodal large language models
Semantic consensus
Cross-modal alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sen Nie
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, and the University of Chinese Academy of Sciences
J
Jie Zhang
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, and the University of Chinese Academy of Sciences
Zhongqi Wang
Zhongqi Wang
Institute of Computing Technology, Chinese Academy of Sciences
Model Robustness
Shiguang Shan
Shiguang Shan
Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition
Xilin Chen
Xilin Chen
Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine Learning