MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of security evaluations against image-carrier attacks in existing multimodal agents, where malicious instructions can be disguised as visual guidance to trigger unauthorized operations. To this end, this work proposes the Native Context Visual Attack (NCVA) paradigm, which embeds malicious instructions into native components of instructional images. Furthermore, it constructs the first end-to-end security benchmark, leveraging isolated sandboxes and automated attack generation techniques for comprehensive evaluation. Across nine model configurations, NCVA consistently induces unauthorized operations with an average attack success rate of 43.1%, significantly outperforming text-based baselines. This research fills a critical gap in runtime visual security assessment and demonstrates that superior task performance does not inherently guarantee system security.
📝 Abstract
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.
Problem

Research questions and friction points this paper is trying to address.

multimodal agent skills
image-borne attacks
agent safety
visual attack benchmark
skill security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Skills
Visual Attack
Safety Benchmark
Agent Security
NCVA
💼 Related Jobs
No related jobs found.
L
Lingqi Jiang
Zhejiang University
J
Jialuo Chen
Zhejiang University
J
Jianan Ma
Hangzhou Dianzi University
X
Xinhao Deng
Ant Group
X
Xiaohu Du
Ant Group
Sibo Yi
Sibo Yi
Tsinghua University
AI safety
Y
Yuqi Qing
Tsinghua University
Zhenguang Liu
Zhenguang Liu
Zhejiang University
BlockchainSmart Contract SecurityMultimedia
Qinming He
Qinming He
Zhejiang University, Professor
BlockchainData miningMachine learning
S
Shiwen Cui
Ant Group
C
Changhua Meng
Ant Group