Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating adversarial robustness in Vision-Language-Action (VLA) models when white-box access is unavailable, proposing a purely black-box attack paradigm leveraging ancestral Vision-Language Models (VLMs). By optimizing adversarial patches on the ancestral VLM, the method disrupts the target model's visual perception and instruction grounding capabilities. To enhance attack transferability, it introduces visual token perturbation, semantic evidence suppression, and a two-stage curriculum learning strategy. This work is the first to reveal how VLAs inherit adversarial vulnerabilities from their ancestral VLMs, significantly degrading task success rates. Furthermore, it delineates the boundaries of vulnerability inheritance, thereby deepening the understanding of security risks within embodied intelligence systems.
📝 Abstract
Vision-Language-Action models (VLAs) are increasingly deployed in safety-critical physical environments, yet their adversarial robustness remains poorly understood. Existing attacks typically assume white-box access or rely on surrogate VLAs, which rarely holds in real-world deployments. Our key insight is that most VLAs are adapted from a publicly released pretrained vision-language model (VLM), inheriting two capabilities essential for action generation: visual perception and instruction-conditioned grounding. Therefore, this paper explores a previously unaddressed question: can an adversary attack deployed VLAs using only their ancestor VLMs? To this end, we propose three adversarial patch attacks that disrupt the inherited capabilities: a vision disruption attack that corrupts the projected visual tokens through relative and absolute terms, an instruction-grounded semantic evidence suppression attack that removes the visual evidence required for instruction-grounded concepts, and a joint attack that unifies both objectives under a two-phase curriculum. Experiments across different VLA families on both simulation and static real-world images show that patches optimized on the ancestor VLM cause substantial degradations in VLA task success rates, demonstrating that VLAs inherit adversarial vulnerabilities alongside their foundational capabilities. This effect is not uniform: it is strongest on tasks that require precise instruction-grounded localization, and nearly vanishes on policies whose adaptation rewrites the shared visual-semantic representation or whose action head iteratively smooths perturbations away. By characterizing the boundary conditions of vulnerability inheritance and providing analysis of why the inheritance effect holds or fails, we advance the understanding of safety for VLA-involved systems.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
adversarial patch attacks
black-box attack
vulnerability inheritance
adversarial robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box adversarial patch attack
Vision-Language-Action models
Ancestor VLM exploitation
Vulnerability inheritance
Instruction-grounded suppression
🔎 Similar Papers
No similar papers found.
X
Xiaoyi Pang
The Hong Kong University of Science and Technology
H
Haoyue Feng
Beihang University
Q
Quanxin Shou
The Hong Kong University of Science and Technology
Y
Yikun Miao
The Hong Kong University of Science and Technology
Z
Zhengyang Yan
The Hong Kong University of Science and Technology
Song Guo
Song Guo
Chair Professor of CSE, HKUST
Large Language ModelEdge AIMachine Learning Systems