Score
Designs, builds, and evaluates mechanisms that detect, analyze, and mitigate prompt‑injection and other malicious or malformed inputs, and that sanitize and validate model outputs to prevent information leakage, unsafe content, or execution of unintended instructions. This work includes creating input/output filters and transformers, parsing and validation pipelines, detection models and heuristics, adversarial tests, and runtime defenses and monitoring to enforce sanitization policies.
Large language models are vulnerable to emerging threats such as prompt injection and jailbreaking attacks, necessitating systematic defense mechanisms. This work addresses this challenge through a comprehensive literature review of 88 relevant studies, extending the NIST adversarial machine learning defense taxonomy by introducing new defense categories and establishing unified terminology. Building upon this refined framework, the study constructs a standardized catalog that evaluates defenses along key dimensions—including effectiveness, open-source availability, and model generality—encompassing mainstream large language models and attack benchmarks. The resulting resource offers researchers and practitioners a reusable, structured guide for implementing robust and interoperable defenses against adversarial exploits in large language models.
This work proposes a systematic security analysis framework for large language models (LLMs) that addresses the complex interdependencies across data, prompts, tool invocations, and other components throughout the full model lifecycle and application stack. Centered on three core failure modes—trust boundary violations, conversion of untrusted data into executable instructions, and authorization escalation errors—the framework spans eight phases from data collection to deployment. Through a structured literature review mapped to key security objectives such as confidentiality, integrity, privacy, and agent control, the study categorizes attack surfaces, evaluation methodologies, and defense mechanisms at each stage. It establishes the first comprehensive vulnerability and defense taxonomy covering the entire LLM stack, identifies critical open challenges, and outlines promising directions including composable security, provenance-aware retrieval, and isolated tool invocation.
This work addresses the vulnerability of retrieval-augmented generation (RAG) and tool-augmented large language models to malicious instructions embedded in external text, which can trigger harmful behaviors. Existing defense mechanisms suffer from poor generalization and susceptibility to optimization-based attacks. To overcome these limitations, the authors propose SONAR, a novel framework that integrates sentence-level relational graphs with natural language inference (NLI). By leveraging entailment and contradiction scores to detect malicious content and applying a connectivity-driven pruning strategy, SONAR achieves effective instruction sanitization without requiring any model retraining. Evaluated across multiple models and datasets, the method reduces attack success rates to near zero and substantially outperforms nine state-of-the-art baseline defenses.
Prompt injection attacks pose a critical security threat to tool-augmented large language model (LLM) agents operating in untrusted environments. To address this, we propose SIC—a novel iterative soft instruction cleansing mechanism. SIC dynamically detects and rectifies missed malicious instructions through multiple rounds of detection and semantic rewriting, overcoming the limitations of single-shot purification. It introduces three key innovations: (1) instruction residue detection to identify residual adversarial content, (2) a maximum iteration bound to ensure computational efficiency, and (3) a safety interruption policy to halt execution upon persistent threats—thereby balancing robustness and controllability. Extensive experiments demonstrate that SIC significantly outperforms existing single-pass methods: under worst-case conditions with strong adversaries, it reduces attack success rates to just 15%, substantially raising the bar for successful exploitation. This work establishes a scalable, lightweight, and practical defense paradigm for securing LLM agents in open, real-world environments.
This study addresses the vulnerability of large language models to prompt injection attacks when sensitive information is embedded in system prompts, which can lead to unintended secret disclosure. The authors propose an adaptive adversarial attack framework that dynamically evolves attack strategies over more than 20,000 red-team evaluations to systematically assess nine state-of-the-art defense mechanisms. Experimental results demonstrate that all defenses relying solely on the model’s intrinsic safeguards fail to prevent information leakage, whereas only application-layer output filtering combined with hard-coded rules achieves zero leakage. This work provides the first large-scale empirical evidence—through adaptive adversarial testing—that security boundaries must be enforced by application code rather than by the model itself.
To address the high runtime overhead imposed by instrumentation-based sanitizers in fuzzing, this paper proposes a lightweight, on-demand detection framework that decouples taint analysis from the fuzzing loop. Methodologically, it introduces (1) a novel execution-mode analysis to precisely identify inputs with potential vulnerability-triggering behavior; (2) dynamic deferred scheduling of sanitizer invocations—enabling sanitizer-augmented builds only for “interesting” inputs; and (3) a synergistic integration of lightweight execution trace capture, pattern matching, and conditional triggering, ensuring compatibility with multiple sanitizers including ASan and UBSan. Implemented atop AFL++, the framework demonstrates superior vulnerability discovery—outperforming all baseline fuzzers across 12 real-world programs within 24 hours—while achieving zero missed detections on known bugs. Crucially, it reduces average overhead by several orders of magnitude compared to conventional sanitizer-integrated fuzzing.
Existing research lacks systematic modeling and standardized evaluation of prompt injection attacks and defenses in LLM-based integrated applications. Method: We propose the first general formal attack framework that unifies five known attack categories and enables derivation of novel composite attacks; we further construct the first open-source, cross-model (e.g., GPT, Llama, Claude) and cross-task (e.g., QA, summarization, reasoning—seven tasks total) benchmark, covering ten defense mechanisms. Contribution/Results: Through red-team/blue-team adversarial experiments, we demonstrate that most existing defenses fail under complex, realistic scenarios. We release Open-Prompt-Injection—a reproducible, multidimensional quantitative evaluation platform—to advance standardization and community collaboration in prompt security research.
This study addresses the vulnerability of AI-powered software reverse engineering agents to prompt injection attacks by presenting the first systematic investigation of adversarial prompt injections embedded within executable binaries and their obfuscated variants. The work proposes an integrated defense framework that combines static analysis, specialized detection algorithms, and deobfuscation techniques to effectively identify diverse prompt injection attacks in decompiled output. Experimental results demonstrate that the proposed approach maintains high detection accuracy even against heavily obfuscated code, significantly enhancing the security and robustness of AI-driven reverse engineering systems in real-world operational environments.
This work systematically investigates the vulnerability of large language models (LLMs) to prompt injection attacks when processing untrusted inputs and evaluates a novel defense strategy that encapsulates such inputs as simulated tool calls to leverage trust isolation mechanisms inherent in the model’s instruction hierarchy. Using an automated red-teaming framework, the authors assess this approach across seven prominent LLMs on three LLM-as-a-Judge tasks. Contrary to expectations, tool encapsulation does not consistently improve robustness; in binary judgment tasks such as GSM8K scoring, it significantly increases attack success rates. Moreover, certain models exhibit instruction hierarchy inversion, wherein higher-level directives are overridden by lower-level injected content. These findings reveal critical limitations and counterintuitive behaviors in current LLM architectures under real-world deployment scenarios.
This work addresses the lack of a unified analytical framework for prompt injection attacks, which are typically represented as unstructured strings, hindering systematic annotation, comparison, and evolution. The authors propose the first structured seven-component model—comprising carrier, delivery vector, obfuscation mechanism, context boundary breach, privilege escalation, payload, and exfiltration channel—that focuses on attacker intent rather than surface-level text to establish a reusable attack parsing framework. This model integrates existing techniques, aligns with Cyber Threat Intelligence (CTI) standards, and enables attack flow graph modeling. Notably, minimal jailbreaking is formalized as a subspace within this framework. Empirical validation on EchoLeak (CVE-2025-32711) and real-world AI evasion malware demonstrates its effectiveness in systematically describing and reproducing prompt injection attacks.
This study addresses the vulnerability of large language models (LLMs) to adversarial prompt injection attacks in the context of Security Operations Center (SOC) log analysis, where logs containing clear indicators of compromise are misclassified as benign. It presents the first systematic investigation into such attacks targeting LLM-based log interpretation, introducing an evaluation framework that generates and optimizes adversarial log samples to assess the robustness of state-of-the-art LLMs. The findings reveal that despite the susceptibility of multiple advanced models to these attacks, their generated explanations often contain subtle yet discernible traces of the injected prompts. These latent cues can be leveraged to effectively detect prompt injection attempts, offering a novel avenue for developing defensive mechanisms against such threats in security-sensitive applications.
This work addresses the vulnerability of large code models to indirect prompt injection attacks concealed in external contexts such as comments and string literals. To mitigate this threat, the authors propose CodeSentinel, a novel three-tier defense framework that integrates syntax-guided pre-filtering, concrete syntax tree (CST)-guided dynamic Min-K% scoring, and node perturbation analysis to accurately detect and neutralize semantic triggers. Leveraging Tree-sitter to extract high-risk CST nodes, CodeSentinel synergistically combines syntactic parsing, dynamic scoring, and perturbation-based detection. Evaluated against six state-of-the-art attack variants, the method achieves an average node-level F1 score of 0.80, substantially outperforming existing baselines including CodeGarrison, DePA, and KillBadCode.