AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "risk-rejection" gap in large audio language models, wherein internal risk recognition fails to translate into refusal behaviors, rendering such models vulnerable to jailbreak attacks. To tackle this, we are the first to formally identify this phenomenon and propose a post-detection intervention defense mechanism. Specifically, our approach leverages hierarchical probing techniques to localize risk representations and employs intermediate-layer gating to dynamically activate safety adapters for selective intervention. This work establishes internal intervention as a novel pathway for enhancing model robustness. Extensive evaluations across multiple benchmarks demonstrate that our method reduces the average unsafe response rate from 17.9% to 0.4%, while introducing only a marginal increase in over-refusal on benign inputs.
📝 Abstract
Large audio-language models (LALMs) expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing reveals the latter: risk-related information remains decodable from intermediate representations, yet the internal risk signal fails to translate into refusal in later-layer processing. We identify this discrepancy as the risk-to-refusal gap. Building on this finding, we propose AEGIS, a detect-then-intervene defense whose mid-layer risk gate selectively activates downstream safety adapters. Across six LALMs and three heterogeneous audio jailbreak benchmarks, AEGIS reduces the average unsafe rate from 17.9% to 0.4%, while causing only a marginal increase in over-refusal on benign inputs. These results establish selective internal intervention as an effective path toward more robust refusal in LALMs. The code is available at https://github.com/azzzzliao/aegis-audio-defense.
Problem

Research questions and friction points this paper is trying to address.

Large Audio-Language Models
Audio Jailbreaks
Risk-to-Refusal Gap
Safety Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Jailbreak Defense
Risk-to-Refusal Gap
Layer-wise Probing
Selective Internal Intervention
Large Audio-Language Models
Y
Yu-Ling Liao
National Taiwan University
T
Tzu-Chin Chiu
National Taiwan University
Z
Zong-You Chen
National Taiwan University
C
Chi-Lei Tsai
National Taiwan University
Shao-Yuan Lo
Shao-Yuan Lo
Assistant Professor, National Taiwan University
Trustworthy AIMachine LearningMultimodal Models