🤖 AI Summary
This study addresses the lack of security evaluations against image-carrier attacks in existing multimodal agents, where malicious instructions can be disguised as visual guidance to trigger unauthorized operations. To this end, this work proposes the Native Context Visual Attack (NCVA) paradigm, which embeds malicious instructions into native components of instructional images. Furthermore, it constructs the first end-to-end security benchmark, leveraging isolated sandboxes and automated attack generation techniques for comprehensive evaluation. Across nine model configurations, NCVA consistently induces unauthorized operations with an average attack success rate of 43.1%, significantly outperforming text-based baselines. This research fills a critical gap in runtime visual security assessment and demonstrates that superior task performance does not inherently guarantee system security.
📝 Abstract
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.