sanitize inputs and outputs

Designs, builds, and evaluates mechanisms that detect, analyze, and mitigate prompt‑injection and other malicious or malformed inputs, and that sanitize and validate model outputs to prevent information leakage, unsafe content, or execution of unintended instructions. This work includes creating input/output filters and transformers, parsing and validation pipelines, detection models and heuristics, adversarial tests, and runtime defenses and monitoring to enforce sanitization policies.

sanitizeinputsandoutputs

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work proposes a systematic security analysis framework for large language models (LLMs) that addresses the complex interdependencies across data, prompts, tool invocations, and other components throughout the full model lifecycle and application stack. Centered on three core failure modes—trust boundary violations, conversion of untrusted data into executable instructions, and authorization escalation errors—the framework spans eight phases from data collection to deployment. Through a structured literature review mapped to key security objectives such as confidentiality, integrity, privacy, and agent control, the study categorizes attack surfaces, evaluation methodologies, and defense mechanisms at each stage. It establishes the first comprehensive vulnerability and defense taxonomy covering the entire LLM stack, identifies critical open challenges, and outlines promising directions including composable security, provenance-aware retrieval, and isolated tool invocation.

Application StackLarge Language ModelLifecycle

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the vulnerability of retrieval-augmented generation (RAG) and tool-augmented large language models to malicious instructions embedded in external text, which can trigger harmful behaviors. Existing defense mechanisms suffer from poor generalization and susceptibility to optimization-based attacks. To overcome these limitations, the authors propose SONAR, a novel framework that integrates sentence-level relational graphs with natural language inference (NLI). By leveraging entailment and contradiction scores to detect malicious content and applying a connectivity-driven pruning strategy, SONAR achieves effective instruction sanitization without requiring any model retraining. Evaluated across multiple models and datasets, the method reduces attack success rates to near zero and substantially outperforms nine state-of-the-art baseline defenses.

adversarial attacksLLM agentsmalicious instructions

Soft Instruction De-escalation Defense

Oct 23, 2025
NP
Nils Philipp Walter
🏛️ CISPA Helmholtz Center for Information Security | Google DeepMind | AI Sequrity Company

Prompt injection attacks pose a critical security threat to tool-augmented large language model (LLM) agents operating in untrusted environments. To address this, we propose SIC—a novel iterative soft instruction cleansing mechanism. SIC dynamically detects and rectifies missed malicious instructions through multiple rounds of detection and semantic rewriting, overcoming the limitations of single-shot purification. It introduces three key innovations: (1) instruction residue detection to identify residual adversarial content, (2) a maximum iteration bound to ensure computational efficiency, and (3) a safety interruption policy to halt execution upon persistent threats—thereby balancing robustness and controllability. Extensive experiments demonstrate that SIC significantly outperforms existing single-pass methods: under worst-case conditions with strong adversaries, it reduces attack success rates to just 15%, substantially raising the bar for successful exploitation. This work establishes a scalable, lightweight, and practical defense paradigm for securing LLM agents in open, real-world environments.

Defending LLM agents from malicious prompt injections in untrusted dataPreventing compromised agent behavior through iterative security checksSanitizing inputs by detecting and rewriting harmful instruction content

This study addresses the vulnerability of large language models to prompt injection attacks when sensitive information is embedded in system prompts, which can lead to unintended secret disclosure. The authors propose an adaptive adversarial attack framework that dynamically evolves attack strategies over more than 20,000 red-team evaluations to systematically assess nine state-of-the-art defense mechanisms. Experimental results demonstrate that all defenses relying solely on the model’s intrinsic safeguards fail to prevent information leakage, whereas only application-layer output filtering combined with hard-coded rules achieves zero leakage. This work provides the first large-scale empirical evidence—through adaptive adversarial testing—that security boundaries must be enforced by application code rather than by the model itself.

Defense EvaluationLarge Language ModelsPrompt Injection

SAND: Decoupling Sanitization from Fuzzing for Low Overhead

Feb 26, 2024
ZK
Ziqiao Kong
🏛️ ETH Zurich | Nanyang Technological University | City University of Hong Kong

To address the high runtime overhead imposed by instrumentation-based sanitizers in fuzzing, this paper proposes a lightweight, on-demand detection framework that decouples taint analysis from the fuzzing loop. Methodologically, it introduces (1) a novel execution-mode analysis to precisely identify inputs with potential vulnerability-triggering behavior; (2) dynamic deferred scheduling of sanitizer invocations—enabling sanitizer-augmented builds only for “interesting” inputs; and (3) a synergistic integration of lightweight execution trace capture, pattern matching, and conditional triggering, ensuring compatibility with multiple sanitizers including ASan and UBSan. Implemented atop AFL++, the framework demonstrates superior vulnerability discovery—outperforming all baseline fuzzers across 12 real-world programs within 24 hours—while achieving zero missed detections on known bugs. Crucially, it reduces average overhead by several orders of magnitude compared to conventional sanitizer-integrated fuzzing.

Disinfectant UsageResource ConsumptionSoftware Testing

Formalizing and Benchmarking Prompt Injection Attacks and Defenses

Oct 19, 2023
YL
Yupei Liu
🏛️ The Pennsylvania State University | Duke University

Existing research lacks systematic modeling and standardized evaluation of prompt injection attacks and defenses in LLM-based integrated applications. Method: We propose the first general formal attack framework that unifies five known attack categories and enables derivation of novel composite attacks; we further construct the first open-source, cross-model (e.g., GPT, Llama, Claude) and cross-task (e.g., QA, summarization, reasoning—seven tasks total) benchmark, covering ten defense mechanisms. Contribution/Results: Through red-team/blue-team adversarial experiments, we demonstrate that most existing defenses fail under complex, realistic scenarios. We release Open-Prompt-Injection—a reproducible, multidimensional quantitative evaluation platform—to advance standardization and community collaboration in prompt security research.

Benchmarking existing attacks and defenses across multiple modelsEstablishing common evaluation framework for future research developmentFormalizing prompt injection attacks in LLM applications systematically

Latest Papers

What's happening recently
View more

This study addresses the vulnerability of AI-powered software reverse engineering agents to prompt injection attacks by presenting the first systematic investigation of adversarial prompt injections embedded within executable binaries and their obfuscated variants. The work proposes an integrated defense framework that combines static analysis, specialized detection algorithms, and deobfuscation techniques to effectively identify diverse prompt injection attacks in decompiled output. Experimental results demonstrate that the proposed approach maintains high detection accuracy even against heavily obfuscated code, significantly enhancing the security and robustness of AI-driven reverse engineering systems in real-world operational environments.

adversarial attacksAI agentsprompt injection

This work systematically investigates the vulnerability of large language models (LLMs) to prompt injection attacks when processing untrusted inputs and evaluates a novel defense strategy that encapsulates such inputs as simulated tool calls to leverage trust isolation mechanisms inherent in the model’s instruction hierarchy. Using an automated red-teaming framework, the authors assess this approach across seven prominent LLMs on three LLM-as-a-Judge tasks. Contrary to expectations, tool encapsulation does not consistently improve robustness; in binary judgment tasks such as GSM8K scoring, it significantly increases attack success rates. Moreover, certain models exhibit instruction hierarchy inversion, wherein higher-level directives are overridden by lower-level injected content. These findings reveal critical limitations and counterintuitive behaviors in current LLM architectures under real-world deployment scenarios.

adversarial robustnessinstruction hierarchylarge language models

This work addresses the lack of a unified analytical framework for prompt injection attacks, which are typically represented as unstructured strings, hindering systematic annotation, comparison, and evolution. The authors propose the first structured seven-component model—comprising carrier, delivery vector, obfuscation mechanism, context boundary breach, privilege escalation, payload, and exfiltration channel—that focuses on attacker intent rather than surface-level text to establish a reusable attack parsing framework. This model integrates existing techniques, aligns with Cyber Threat Intelligence (CTI) standards, and enables attack flow graph modeling. Notably, minimal jailbreaking is formalized as a subspace within this framework. Empirical validation on EchoLeak (CVE-2025-32711) and real-world AI evasion malware demonstrates its effectiveness in systematically describing and reproducing prompt injection attacks.

attack labelingcyber threat intelligencelarge language models

This study addresses the vulnerability of large language models (LLMs) to adversarial prompt injection attacks in the context of Security Operations Center (SOC) log analysis, where logs containing clear indicators of compromise are misclassified as benign. It presents the first systematic investigation into such attacks targeting LLM-based log interpretation, introducing an evaluation framework that generates and optimizes adversarial log samples to assess the robustness of state-of-the-art LLMs. The findings reveal that despite the susceptibility of multiple advanced models to these attacks, their generated explanations often contain subtle yet discernible traces of the injected prompts. These latent cues can be leveraged to effectively detect prompt injection attempts, offering a novel avenue for developing defensive mechanisms against such threats in security-sensitive applications.

adversarial log injectionLLM-based log interpretationprompt injection

This work addresses the vulnerability of large code models to indirect prompt injection attacks concealed in external contexts such as comments and string literals. To mitigate this threat, the authors propose CodeSentinel, a novel three-tier defense framework that integrates syntax-guided pre-filtering, concrete syntax tree (CST)-guided dynamic Min-K% scoring, and node perturbation analysis to accurately detect and neutralize semantic triggers. Leveraging Tree-sitter to extract high-risk CST nodes, CodeSentinel synergistically combines syntactic parsing, dynamic scoring, and perturbation-based detection. Evaluated against six state-of-the-art attack variants, the method achieves an average node-level F1 score of 0.80, substantially outperforming existing baselines including CodeGarrison, DePA, and KillBadCode.

adversarial attackscode contextcode large language models

Hot Scholars

AH

Amir Houmansadr

University of Massachusetts Amherst
Privacy-enhancing technologiesTrustworthy MLNetwork traffic analysis
PM

Prateek Mittal

Professor, Princeton University
Security and PrivacySystems and NetworkingMachine Learning
JJ

Jinyuan Jia

Assistant Professor, Penn State
AI Security
CS

Chawin Sitawarin

Google DeepMind
machine learningartificial intelligencesecurityprivacy
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design