Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过XAI技术分析Prompt Guard 2如何区分恶意与良性提示,揭示了其决策机制,并探讨了对抗性攻击的可能性及设计改进。
📝 Abstract
Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their internal decision logic is largely opaque to both defenders and attackers. This paper presents an exploratory case study that applies explainable artificial intelligence (XAI) techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts. We conduct four experiments to probe this question empirically. Guided by Vanilla Gradient and SHAP attributions, we find that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substitution and sentence-level paraphrasing can flip its predictions while altering only a moderate fraction of the text, in some cases yielding a successful jailbreak against the underlying LLM. A dataset-scale saliency analysis further shows that undetected injection prompts systematically lack the lexical markers the classifier relies on. We discuss the implications of these findings for the design and evaluation of classifier-based guardrails, and argue that explanation methods intended to support transparency can simultaneously lower the cost of constructing successful adversarial bypasses.
Problem

Research questions and friction points this paper is trying to address.

Large language models
Adversarial manipulation
Prompt injection
Jailbreak attacks
Classifier-based guardrails
Innovation

Methods, ideas, or system contributions that make the work stand out.

XAI
SHAP
Saliency Analysis
Adversarial Bypass
🔎 Similar Papers
No similar papers found.
F
Fernando Outeda
Pedeciba Informática, Uruguay
G
Gustavo Betarte
Pedeciba Informática, Uruguay; InCo, Facultad de Ingeniería, Universidad de la República
J
Juan Diego Campo
Pedeciba Informática, Uruguay; InCo, Facultad de Ingeniería, Universidad de la República
Fiorella Cravero
Fiorella Cravero
Departamento de Informática e Inteligencia Artificial, Universidad Católica del Uruguay