A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy

📅 2026-01-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models are vulnerable to emerging threats such as prompt injection and jailbreaking attacks, necessitating systematic defense mechanisms. This work addresses this challenge through a comprehensive literature review of 88 relevant studies, extending the NIST adversarial machine learning defense taxonomy by introducing new defense categories and establishing unified terminology. Building upon this refined framework, the study constructs a standardized catalog that evaluates defenses along key dimensions—including effectiveness, open-source availability, and model generality—encompassing mainstream large language models and attack benchmarks. The resulting resource offers researchers and practitioners a reusable, structured guide for implementing robust and interoperable defenses against adversarial exploits in large language models.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Adversarial Attacks & Robustness

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The rapid advancement and widespread adoption of generative artificial intelligence (GenAI) and large language models (LLMs) has been accompanied by the emergence of new security vulnerabilities and challenges, such as jailbreaking and other prompt injection attacks. These maliciously crafted inputs can exploit LLMs, causing data leaks, unauthorized actions, or compromised outputs, for instance. As both offensive and defensive prompt injection techniques evolve quickly, a structured understanding of mitigation strategies becomes increasingly important. To address that, this work presents the first systematic literature review on prompt injection mitigation strategies, comprehending 88 studies. Building upon NIST's report on adversarial machine learning, this work contributes to the field through several avenues. First, it identifies studies beyond those documented in NIST's report and other academic reviews and surveys. Second, we propose an extension to NIST taxonomy by introducing additional categories of defenses. Third, by adopting NIST's established terminology and taxonomy as a foundation, we promote consistency and enable future researchers to build upon the standardized taxonomy proposed in this work. Finally, we provide a comprehensive catalog of the reviewed prompt injection defenses, documenting their reported quantitative effectiveness across specific LLMs and attack datasets, while also indicating which solutions are open-source and model-agnostic. This catalog, together with the guidelines presented herein, aims to serve as a practical resource for researchers advancing the field of adversarial machine learning and for developers seeking to implement effective defenses in production systems.
Problem

Research questions and friction points this paper is trying to address.

prompt injection
jailbreaking
large language models
security vulnerabilities
adversarial attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt injection
jailbreaking
LLM defenses
NIST taxonomy extension
systematic literature review
P
Pedro H. Barcha Correia
Universidade de São Paulo (USP), Laboratório de Arquitetura e Redes de Computadores (LARC), Brazil
R
R. W. Achjian
Universidade de São Paulo (USP), Laboratório de Arquitetura e Redes de Computadores (LARC), Brazil
D
Diego E. G. Caetano de Oliveira
Universidade do Estado de Santa Catarina (UDESC), Programa de Pós-Graduação em Computação Aplicada (PPGCAP), Brazil
Y
Ygor Acacio Maria
Universidade de São Paulo (USP), Laboratório de Arquitetura e Redes de Computadores (LARC), Brazil
V
Victor Takashi Hayashi
Universidade de São Paulo (USP), Laboratório de Arquitetura e Redes de Computadores (LARC), Brazil
Marcos Lopes
Marcos Lopes
Professor of Linguistics, Universidade de São Paulo
SemanticsComputational LinguisticCognitive Science
C
C. Miers
Universidade do Estado de Santa Catarina (UDESC), Programa de Pós-Graduação em Computação Aplicada (PPGCAP), Brazil
M
Marcos A. Simplício
Universidade de São Paulo (USP), Laboratório de Arquitetura e Redes de Computadores (LARC), Brazil