Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of multimodal large language model agents to stealthy concurrent audio prompt injection attacks during continuous audio interactions, which jeopardize instruction execution security. The authors propose a highly covert attack method that imperceptibly overlays malicious instructions onto user speech to achieve concurrent injection and introduce AudioAgentSecurity, the first benchmark for evaluating audio-based instruction injection attacks. To counter this threat, they design a novel defense framework, CADV, which integrates sound source separation with cross-modal consistency analysis to detect such attacks. Experiments demonstrate an average attack success rate of 69.10% across 11 mainstream agents, including Gemini 3 Pro, while CADV achieves over 90% detection accuracy. Real-world evaluations on mobile devices confirm both the stealthiness of the attack and the effectiveness of the proposed defense.
📝 Abstract
Large Language Model (LLM)-driven multimodal agents are increasingly deployed to execute autonomous tasks via continuous audio interaction. While this paradigm enhances interaction naturalness, it introduces a critical yet under-explored attack surface, as audio inputs inevitably contain environmental noise beyond user control. In this paper, we investigate concurrent audio prompt injection attacks targeting multimodal agents. Distinct from traditional acoustic attacks on voice devices, we propose novel techniques for instruction augmentation and scenario concealment. These methods allow malicious audio instructions to imperceptibly "piggyback" onto user speech, thereby hijacking agents to execute malicious actions. To systematically quantify this threat, we construct AudioAgentSecurity, the first comprehensive benchmark for audio instruction injection attacks, encompassing 8 real-world task scenarios and 10 distinct attack patterns. We evaluate 11 state-of-the-art agents, including Gemini 3 Pro and GPT-4o-audio. Notably, our methods achieve an average Attack Success Rate (ASR) of 69.10\% against the advanced Gemini 3 Pro. To counter this threat, we further introduce Cascaded Audio Decoupling and Verification (CADV), a defense mechanism based on source separation and consistency analysis. Compared with existing prompt-level defenses, CADV leverages acoustic source separation and cross-modal consistency analysis to detect audio instruction injections more robustly, achieving over 90\% detection success across diverse attack vectors. Finally, real-world experiments with human volunteers on Doubao AI Smartphone in diverse dynamic real-world scenarios confirm the attacks' high stealth and efficacy, while demonstrating that our defense reliably mitigates these vulnerabilities.
Problem

Research questions and friction points this paper is trying to address.

audio prompt injection
multimodal LLM agents
stealthy attack
concurrent audio
adversarial audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio prompt injection
multimodal LLM agents
stealthy attack
source separation
cross-modal consistency
🔎 Similar Papers
No similar papers found.
M
Mingxiao Liu
Hangzhou Dianzi University
Y
Yitong Li
Hangzhou Dianzi University
H
Haoren Zhao
Hangzhou Dianzi University
Y
Yaoxiang Bian
Hangzhou Dianzi University
J
Jianan Ma
Hangzhou Dianzi University, Ant Group
Jian Zhang
Jian Zhang
Hangzhou Dianzi University
Dynamic networkGraph neural networkAnomaly detection
J
Jialuo Chen
Zhejiang University, Ant Group
X
Xinhao Deng
Tsinghua University, Ant Group
Z
Zhen Wang
Hangzhou Dianzi University