adversarial prompt generation

Designs, builds, or analyzes methods that automatically generate or optimize prompts to induce incorrect, unwanted, or targeted behaviors in large language models, including mutations of seed prompts, minimal token edits, jailbreaks, and prompt rewrites. Work covers algorithmic red‑teaming and policy learning, black‑box and meta/targeted prompt optimization across tasks and languages, practical operation via standard API access, and quantitative measurement of attack success while separating adversarial intent from toxicity.

adversarialpromptgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs

Jul 13, 2025
MD
MohammadReza Davari
🏛️ Concordia University | Mila – Quebec AI Institute | Microsoft

This work addresses two key limitations in black-box large language model (LLM) prompt optimization: underutilization of correct prediction signals and poor cross-model transferability. To this end, we propose an enhanced feedback-driven prompt optimization framework. Methodologically, it introduces a dual-track reinforcement mechanism—retaining effective prompt components via both positive and negative signals—integrates text-gradient reconstruction, multi-signal feedback aggregation, and noise filtering, and incorporates an explicit prompt transfer strategy. Our key contribution is the first systematic integration of positive reinforcement learning into automated prompt optimization, enabling active exploitation of correct prediction information. Experiments demonstrate that our method consistently outperforms strong baselines on both standard prompt optimization and cross-model/cross-API transfer tasks, achieving simultaneous improvements in accuracy, convergence speed, and computational efficiency.

Enabling efficient prompt migration across different LLM versionsOptimizing prompts using both positive and negative reinforcement signalsReducing noise in LLM feedback via diversification techniques

This study investigates how prompt quality affects the security of code generated by large language models (LLMs). We identify that “benign but poorly formulated” prompts significantly increase the likelihood of security vulnerabilities in generated code. To address this, we propose a three-dimensional evaluation framework—assessing goal clarity, information completeness, and logical consistency—and introduce CWE-BENCH-PYTHON, the first normative, graded benchmark dataset for Python focused on Common Weakness Enumeration (CWE) vulnerabilities. Our empirical analysis is the first to systematically demonstrate a strong negative correlation between prompt normativity and the incidence of CWE-class security defects: lower prompt quality strongly predicts higher vulnerability rates. Furthermore, we integrate advanced prompting techniques—including Chain-of-Thought reasoning and Self-Correction—to substantially reduce unsafe code generation. Collectively, these findings establish “improving user prompt quality” as a novel paradigm for enhancing the security of AI-generated code, providing both theoretical foundations and practical methodologies for security-aware prompt engineering.

Developing mitigation strategies through advanced prompting techniques for code safetyEstablishing correlation between prompt normativity and insecure code generation ratesInvestigating how poor prompt quality induces security defects in AI-generated code

Large language models are vulnerable to prompt-based attacks (jailbreaking) that circumvent safety mechanisms and elicit harmful content. This work proposes the THREAT framework, which formalizes adversarial prompt generation as a non-convex optimization problem for the first time. By integrating multi-agent collaborative reasoning, iterative adversarial search, and language model redirection techniques, THREAT efficiently produces highly stealthy jailbreaking prompts. Experimental results demonstrate that the method significantly outperforms existing attack strategies across multiple models and datasets, achieving higher attack success rates with lower computational overhead. Notably, fewer than 1% of the generated prompts are flagged as harmful—despite an original refusal rate of approximately 50%—thereby exposing previously undetected security vulnerabilities in current alignment approaches.

adversarial promptsharmful generationjailbreaking

Prompt Optimization and Evaluation for LLM Automated Red Teaming

Jul 29, 2025
MF
Michael Freenor
🏛️ Fuel iX Applied Research | North Carolina State University | University of Minnesota | TELUS Digital | Learn Prompting

This paper addresses the low efficiency and unstable vulnerability discovery of large language models (LLMs) in automated red-teaming, particularly in attack prompt generation. We propose a discoverability-based prompt optimization method that quantifies exploitability by estimating the expected success rate of a single attack on the target system via multi-random-seed sampling; this metric then guides iterative prompt refinement. Attack Success Rate (ASR) serves as the core evaluation metric, enhanced by target-environment randomization and repeated sampling to improve assessment robustness. Our key contribution is the first formal modeling of discoverability as a measurable, optimizable prompt quality metric—eliminating reliance on manual annotations or fixed benchmarks. Experiments demonstrate significant improvements in vulnerability identification rate and cross-model generalization, thereby enhancing both the effectiveness and stability of automated red-teaming.

Improving robustness of automated red teaming evaluationMeasuring attack discoverability via repeated executionsOptimizing prompts for LLM-based attack generation

Prompting Techniques for Secure Code Generation: A Systematic Investigation

Jul 09, 2024
CT
Catherine Tony
🏛️ Hamburg University of Technology

This study systematically investigates how prompt engineering techniques affect the security of code generated by large language models (LLMs). Addressing the problem of security vulnerabilities in LLM-generated code, we evaluate diverse prompting strategies on GPT-3, GPT-3.5, and GPT-4 using 150 security-sensitive natural language instructions. We propose the first systematic taxonomy of prompt engineering tailored to secure code generation. Our key method, Recursive Criticism and Improvement (RCI), iteratively refines code through security-focused critique and revision. Results show RCI consistently reduces security defect rates across all models—by an average of 37.2%—and improves OWASP Top 10 vulnerability detection accuracy by 19.6% over baseline prompting, without compromising functional correctness. This work establishes the first empirically grounded, prompt-based framework for enhancing code security in LLMs, accompanied by a reproducible, practice-oriented guideline for secure AI-assisted development.

Impact of prompting techniquesReduction in security weaknessesSecure code generation by LLMs

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing automated red-teaming approaches, which are confined to optimizing fixed, human-crafted prompts and struggle to dynamically evolve attack strategies. The authors propose an agent-driven, program-level search framework that iteratively edits executable attack programs—rather than single prompts—in a gradient-free, black-box setting, leveraging evaluation feedback to automatically evolve the structural composition of attack strategies. This approach uniquely enables expansion of attack components and modification of control flow, transcending the expressivity constraints of conventional prompt optimization, all without requiring model fine-tuning, human annotation, or GPU resources. Experiments demonstrate an average 17.0-percentage-point improvement in attack success rate across eleven mainstream models, with gains of up to 16 percentage points on state-of-the-art models, underscoring the effectiveness and generality of program-level strategy search.

attack strategyjailbreaklarge language models

This work investigates the differential robustness of large language models (LLMs) aligned via supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning from human feedback (RLHF) under prompt-based adversarial attacks, focusing on attack success rate (ASR) sensitivity to subtle prompt perturbations. Method: We propose a systematic evaluation framework grounded in statistical hypothesis testing, overcoming limitations of conventional attack benchmarks by quantifying variability and significance of ASR shifts. Contribution/Results: Experiments reveal that minute prompt modifications—e.g., word reordering or semantically equivalent substitutions—induce substantial ASR fluctuations (±30% or more), with marked heterogeneity across alignment methods: DPO models exhibit heightened sensitivity to syntactic perturbations, whereas RLHF models degrade more readily under semantic equivalence transformations. To our knowledge, this is the first study to quantitatively establish a strong coupling between alignment methodology and prompt robustness, providing both theoretical insight and empirical grounding for developing interference-resilient, trustworthy alignment evaluation protocols.

Analyzing how prompt perturbations affect attack success rates in LLMsEvaluating vulnerability differences across SFT, DPO and RLHF alignment methodsInvestigating limitations of existing attack benchmarks through statistical analysis

Although large language models (LLMs) exhibit remarkable capabilities, they remain vulnerable to adversarial prompt attacks that can circumvent alignment-based defenses. This work presents the first formal game-theoretic model of the strategic interaction between attackers and defenders in this context, revealing an inherent advantage for the attacker. By integrating game theory, adversarial prompt modeling, and equilibrium analysis, the study derives theoretically provable optimal strategies for both attack and defense. Experimental results across multiple mainstream LLMs and benchmark datasets demonstrate that the proposed theoretically optimal attack strategy substantially outperforms existing methods, offering a rigorous theoretical foundation and practical guidance for the secure deployment of LLMs.

adversarial promptingAI safetyalignment

This work addresses the strong dependence of large language models’ code generation performance on prompt quality and the lack of efficient automated prompt optimization methods. The authors formulate prompt optimization as a sequential decision-making problem and propose a reinforcement learning framework that iteratively refines prompts through a hybrid action space comprising direct generation, genetic-style mutation, and semantic rewriting. A shaping reward mechanism based on unit test feedback guides the optimization process. To the best of our knowledge, this is the first approach to combine reinforcement learning with a hybrid action space for prompt optimization. Using the PPO algorithm, the method is trained on frozen-weight models—including CodeT5+, CodeLLaMA, and DeepSeek-Coder—and achieves substantial improvements over baselines such as EPiC and Reflexion on MBPP+, HumanEval+, and APPS benchmarks, attaining Strict Pass@1 scores of 57.58%–85.50% and Soft-Pass@1 scores of 67.90%–88.20% on MBPP+.

Functional CorrectnessLLM Code GenerationPrompt Optimization

Existing prompt engineering lacks formal determinism guarantees, hindering reliable deployment of large language models (LLMs) in safety-critical applications. Method: We propose a programmable, self-optimizing LLM orchestration framework built upon a generator-auditor-optimizer tripartite adversarial feedback loop. We introduce the Adversarial Trinity topology—the first to model prompts as differentiable semantic variables—and enable gradient-based robust reasoning driven by textual critique. By unifying DSPy’s declarative programming with TextGrad’s text-based differentiation, we integrate semantic computation graphs with adversarial training. Contribution: We establish “observable software engineering” as a new paradigm; formally prove protocol convergence and collapse resistance; and significantly suppress hallucination, delivering deterministic behavioral guarantees in multi-step complex reasoning tasks.

Formalizes LLM orchestration as a programmable, self-optimizing systemMitigates hallucination and prevents model collapse in LLMsProvides deterministic guarantees for mission-critical LLM applications

Hot Scholars

MB

Michael Backes

Chairman and Founding Director of the CISPA Helmholtz Center for Information Security
SecurityprivacycryptographyAI
CX

Chaowei Xiao

University of Wisconsin - Madison/NVIDIA
Trustworthy Machine LearningAdversarial Machine LearningAI SafetyRobust AI
DL

David Lindner

Google DeepMind
Reinforcement LearningScalable OversightActive LearningInterpretability
SA

Sahar Abdelnabi

AI Security Researcher, Microsoft
AI SecurityAI SafetyAdversarial Machine LearningLLMs
WZ

Wanlei Zhou

Professor, City University of Macau, Macao
Parallel and Distributed SystemsIT SecuritySecurity and PrivacyCyber Security