Score
Designs, builds, or analyzes methods that automatically generate or optimize prompts to induce incorrect, unwanted, or targeted behaviors in large language models, including mutations of seed prompts, minimal token edits, jailbreaks, and prompt rewrites. Work covers algorithmic red‑teaming and policy learning, black‑box and meta/targeted prompt optimization across tasks and languages, practical operation via standard API access, and quantitative measurement of attack success while separating adversarial intent from toxicity.
Frequent jailbreaking attacks against large language models (LLMs) and fragmented evaluation criteria hinder progress in prompt security research. Method: This paper introduces the first systematic framework for prompt security, featuring a multi-level taxonomy of attacks and defenses, formalized threat models and cost assumptions, machine-readable safety evaluation profiles, and an open-source benchmarking toolchain. Contributions: (1) We release JAILBREAKDB—the largest human-verified dataset of jailbreaking and benign prompts to date, containing over 120,000 samples; (2) we establish the first open, reproducible, and auditable standardized evaluation benchmark for prompt security; and (3) we conduct a unified, cross-method assessment and ranking of 56 state-of-the-art attack and defense techniques. The framework significantly enhances comparability and reproducibility across studies, providing foundational infrastructure for rigorous, scalable prompt security research.
This work addresses two key limitations in black-box large language model (LLM) prompt optimization: underutilization of correct prediction signals and poor cross-model transferability. To this end, we propose an enhanced feedback-driven prompt optimization framework. Methodologically, it introduces a dual-track reinforcement mechanism—retaining effective prompt components via both positive and negative signals—integrates text-gradient reconstruction, multi-signal feedback aggregation, and noise filtering, and incorporates an explicit prompt transfer strategy. Our key contribution is the first systematic integration of positive reinforcement learning into automated prompt optimization, enabling active exploitation of correct prediction information. Experiments demonstrate that our method consistently outperforms strong baselines on both standard prompt optimization and cross-model/cross-API transfer tasks, achieving simultaneous improvements in accuracy, convergence speed, and computational efficiency.
This study investigates how prompt quality affects the security of code generated by large language models (LLMs). We identify that “benign but poorly formulated” prompts significantly increase the likelihood of security vulnerabilities in generated code. To address this, we propose a three-dimensional evaluation framework—assessing goal clarity, information completeness, and logical consistency—and introduce CWE-BENCH-PYTHON, the first normative, graded benchmark dataset for Python focused on Common Weakness Enumeration (CWE) vulnerabilities. Our empirical analysis is the first to systematically demonstrate a strong negative correlation between prompt normativity and the incidence of CWE-class security defects: lower prompt quality strongly predicts higher vulnerability rates. Furthermore, we integrate advanced prompting techniques—including Chain-of-Thought reasoning and Self-Correction—to substantially reduce unsafe code generation. Collectively, these findings establish “improving user prompt quality” as a novel paradigm for enhancing the security of AI-generated code, providing both theoretical foundations and practical methodologies for security-aware prompt engineering.
Large language models are vulnerable to prompt-based attacks (jailbreaking) that circumvent safety mechanisms and elicit harmful content. This work proposes the THREAT framework, which formalizes adversarial prompt generation as a non-convex optimization problem for the first time. By integrating multi-agent collaborative reasoning, iterative adversarial search, and language model redirection techniques, THREAT efficiently produces highly stealthy jailbreaking prompts. Experimental results demonstrate that the method significantly outperforms existing attack strategies across multiple models and datasets, achieving higher attack success rates with lower computational overhead. Notably, fewer than 1% of the generated prompts are flagged as harmful—despite an original refusal rate of approximately 50%—thereby exposing previously undetected security vulnerabilities in current alignment approaches.
This paper addresses the low efficiency and unstable vulnerability discovery of large language models (LLMs) in automated red-teaming, particularly in attack prompt generation. We propose a discoverability-based prompt optimization method that quantifies exploitability by estimating the expected success rate of a single attack on the target system via multi-random-seed sampling; this metric then guides iterative prompt refinement. Attack Success Rate (ASR) serves as the core evaluation metric, enhanced by target-environment randomization and repeated sampling to improve assessment robustness. Our key contribution is the first formal modeling of discoverability as a measurable, optimizable prompt quality metric—eliminating reliance on manual annotations or fixed benchmarks. Experiments demonstrate significant improvements in vulnerability identification rate and cross-model generalization, thereby enhancing both the effectiveness and stability of automated red-teaming.
This study systematically investigates how prompt engineering techniques affect the security of code generated by large language models (LLMs). Addressing the problem of security vulnerabilities in LLM-generated code, we evaluate diverse prompting strategies on GPT-3, GPT-3.5, and GPT-4 using 150 security-sensitive natural language instructions. We propose the first systematic taxonomy of prompt engineering tailored to secure code generation. Our key method, Recursive Criticism and Improvement (RCI), iteratively refines code through security-focused critique and revision. Results show RCI consistently reduces security defect rates across all models—by an average of 37.2%—and improves OWASP Top 10 vulnerability detection accuracy by 19.6% over baseline prompting, without compromising functional correctness. This work establishes the first empirically grounded, prompt-based framework for enhancing code security in LLMs, accompanied by a reproducible, practice-oriented guideline for secure AI-assisted development.
This work addresses the limitations of existing automated red-teaming approaches, which are confined to optimizing fixed, human-crafted prompts and struggle to dynamically evolve attack strategies. The authors propose an agent-driven, program-level search framework that iteratively edits executable attack programs—rather than single prompts—in a gradient-free, black-box setting, leveraging evaluation feedback to automatically evolve the structural composition of attack strategies. This approach uniquely enables expansion of attack components and modification of control flow, transcending the expressivity constraints of conventional prompt optimization, all without requiring model fine-tuning, human annotation, or GPU resources. Experiments demonstrate an average 17.0-percentage-point improvement in attack success rate across eleven mainstream models, with gains of up to 16 percentage points on state-of-the-art models, underscoring the effectiveness and generality of program-level strategy search.
This work investigates the differential robustness of large language models (LLMs) aligned via supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning from human feedback (RLHF) under prompt-based adversarial attacks, focusing on attack success rate (ASR) sensitivity to subtle prompt perturbations. Method: We propose a systematic evaluation framework grounded in statistical hypothesis testing, overcoming limitations of conventional attack benchmarks by quantifying variability and significance of ASR shifts. Contribution/Results: Experiments reveal that minute prompt modifications—e.g., word reordering or semantically equivalent substitutions—induce substantial ASR fluctuations (±30% or more), with marked heterogeneity across alignment methods: DPO models exhibit heightened sensitivity to syntactic perturbations, whereas RLHF models degrade more readily under semantic equivalence transformations. To our knowledge, this is the first study to quantitatively establish a strong coupling between alignment methodology and prompt robustness, providing both theoretical insight and empirical grounding for developing interference-resilient, trustworthy alignment evaluation protocols.
Although large language models (LLMs) exhibit remarkable capabilities, they remain vulnerable to adversarial prompt attacks that can circumvent alignment-based defenses. This work presents the first formal game-theoretic model of the strategic interaction between attackers and defenders in this context, revealing an inherent advantage for the attacker. By integrating game theory, adversarial prompt modeling, and equilibrium analysis, the study derives theoretically provable optimal strategies for both attack and defense. Experimental results across multiple mainstream LLMs and benchmark datasets demonstrate that the proposed theoretically optimal attack strategy substantially outperforms existing methods, offering a rigorous theoretical foundation and practical guidance for the secure deployment of LLMs.
This work addresses the strong dependence of large language models’ code generation performance on prompt quality and the lack of efficient automated prompt optimization methods. The authors formulate prompt optimization as a sequential decision-making problem and propose a reinforcement learning framework that iteratively refines prompts through a hybrid action space comprising direct generation, genetic-style mutation, and semantic rewriting. A shaping reward mechanism based on unit test feedback guides the optimization process. To the best of our knowledge, this is the first approach to combine reinforcement learning with a hybrid action space for prompt optimization. Using the PPO algorithm, the method is trained on frozen-weight models—including CodeT5+, CodeLLaMA, and DeepSeek-Coder—and achieves substantial improvements over baselines such as EPiC and Reflexion on MBPP+, HumanEval+, and APPS benchmarks, attaining Strict Pass@1 scores of 57.58%–85.50% and Soft-Pass@1 scores of 67.90%–88.20% on MBPP+.
Existing prompt engineering lacks formal determinism guarantees, hindering reliable deployment of large language models (LLMs) in safety-critical applications. Method: We propose a programmable, self-optimizing LLM orchestration framework built upon a generator-auditor-optimizer tripartite adversarial feedback loop. We introduce the Adversarial Trinity topology—the first to model prompts as differentiable semantic variables—and enable gradient-based robust reasoning driven by textual critique. By unifying DSPy’s declarative programming with TextGrad’s text-based differentiation, we integrate semantic computation graphs with adversarial training. Contribution: We establish “observable software engineering” as a new paradigm; formally prove protocol convergence and collapse resistance; and significantly suppress hallucination, delivering deterministic behavioral guarantees in multi-step complex reasoning tasks.