Score
Designs and implements automated pipelines, tooling, and processes that generate, execute, and analyze adversarial test cases and attack scenarios against a target system or model; these systems produce adversarial inputs at scale (e.g., prompt families), run and compare multiple jailbreak/adversary methods, and surface harmful completions and vulnerabilities. They also quantify and report how well safety mitigations hold up across intent taxonomies and other evaluation axes.
Frequent jailbreaking attacks against large language models (LLMs) and fragmented evaluation criteria hinder progress in prompt security research. Method: This paper introduces the first systematic framework for prompt security, featuring a multi-level taxonomy of attacks and defenses, formalized threat models and cost assumptions, machine-readable safety evaluation profiles, and an open-source benchmarking toolchain. Contributions: (1) We release JAILBREAKDB—the largest human-verified dataset of jailbreaking and benign prompts to date, containing over 120,000 samples; (2) we establish the first open, reproducible, and auditable standardized evaluation benchmark for prompt security; and (3) we conduct a unified, cross-method assessment and ranking of 56 state-of-the-art attack and defense techniques. The framework significantly enhances comparability and reproducibility across studies, providing foundational infrastructure for rigorous, scalable prompt security research.
Traditional defense mechanisms are vulnerable to model-guided automated attacks due to their predictable rejection feedback, enabling high-success-rate jailbreaks and prompt injection threats against AI systems. This work proposes a novel “mislead-after-detection” paradigm that shifts the defensive objective from complete attack blocking to reducing the efficiency of attackers’ strategy optimization by generating safe yet misleading responses to confound adversarial judgment. Leveraging probabilistic modeling, we design a lightweight Contextual Misleading via Probabilistic Estimation (CMPE) mechanism and theoretically prove that it asymptotically bounds attack success rates. Evaluated on standard jailbreaking benchmarks, CMPE reduces the upper bound of attack success rates by up to two orders of magnitude, effectively eliminating nearly all successful end-to-end automated attacks.
Large language models (LLMs) remain vulnerable to narrative-style jailbreaking attacks, undermining their reliability in cybersecurity applications. To address this, we propose Jailbreak Mimicry: an automated red-teaming framework that leverages a compact attack model—LoRA-finetuned Mistral-7B—to learn jailbreaking patterns from AdvBench and generate narrative prompts targeting safety-alignment deficiencies. Our approach elevates jailbreak prompt generation from heuristic design to a reproducible, scientific methodology. It is the first to systematically uncover cross-model vulnerabilities across technical and deceptive task domains. Evaluated on GPT-OSS-20B, Jailbreak Mimicry achieves an 81.0% attack success rate—54× higher than baseline methods—and demonstrates strong transferability to mainstream models including GPT-4 and Llama-3. Robustness is further enhanced through joint automated assessment using Claude Sonnet 4 and human validation.
This work addresses three critical gaps in defending large language models (LLMs) against jailbreaking attacks: fragmented defense methodologies, unsystematic evaluation protocols, and poor out-of-distribution (OOD) generalization. To this end, we introduce the first unified, cross-style and cross-distribution evaluation framework for systematically assessing the robustness of 15 mainstream safety guardrails under diverse prompt injection attacks. Our methodology comprises standardized malicious/benign datasets, a multi-dimensional adversarial prompt benchmark, defense response consistency analysis, and principled OOD generalization metrics. Key findings reveal pervasive attack-style bias across existing guardrails; notably, several simple baseline methods surpass state-of-the-art defenses by 12–28% in accuracy under OOD conditions. This study exposes fundamental limitations in current defense evaluation practices and establishes a reproducible, scalable benchmark—grounded in empirical evidence—to advance robust alignment research.
This study addresses the critical challenges of insufficient test data diversity and low vulnerability detection rates in automated software testing. We conduct the first systematic survey of constraint-based adversarial learning methods tailored for software testing, integrating a structured literature review (SLR) with controllable adversarial perturbation modeling. Our analysis yields a taxonomy of constraint-aware adversarial generation techniques, categorizing five distinct technical pathways. We identify key cross-domain barriers and research gaps impeding the transfer of AI security methodologies to software engineering practice. Furthermore, we propose three actionable directions for enhancing automated testing tools—improving functional specificity, vulnerability-triggering capability, and robustness of generated test inputs. Empirical validation demonstrates significant gains in both test effectiveness and fault revelation. The work establishes a theoretical framework and practical guidelines for developing intelligent, resilient, and high-assurance testing tools.
This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.
This study addresses the limitations of traditional red-teaming evaluations that rely solely on attack success rate (ASR) by introducing process mining to analyze the temporal dynamics of large language models during adversarial interactions. Leveraging 8,575 annotated events, the authors construct direct-follow graphs and state transition matrices to uncover dynamic defense mechanisms. Their analysis reveals that GPT-OSS exhibits strong refusal behavior akin to an absorbing state, whereas Llama models display multiple vulnerable pathways susceptible to exploitation. Furthermore, significant differences emerge across models in terms of mutator efficacy and jailbreak time distributions. By moving beyond static ASR metrics, this approach elucidates structural disparities in defensive strategies among models, offering a more nuanced understanding of their robustness against adversarial attacks.
研究针对大型语言模型的越狱攻击,提出系统化防御组合方法,通过标准化评估框架,在保持模型实用性的同时提高安全性。
研究通过电路发现方法分析LLM的越狱行为,使用边缘归因修补和子网络探测技术识别并消减导致安全绕过的计算电路,降低攻击成功率。
This work addresses the vulnerability of large language models (LLMs) to jailbreak attacks by proposing a novel, lightweight pre-defense mechanism. Existing pre-defense approaches suffer from high false-negative rates due to their reliance solely on user prompts, while post-hoc defenses incur substantial computational overhead. To overcome these limitations, the authors introduce a method that leverages a small language model (SLM) to generate draft responses, which—combined with the original user prompt—are fed into a security detection module. By analyzing the transferability of jailbreak attacks between LLMs and SLMs, the approach integrates speculative inference into safety verification. This strategy significantly improves detection accuracy and reduces both false negatives and computational costs, achieving an efficient and low-latency defense without compromising responsiveness.
为解决LLM代理在产品级执行中因越狱引发的风险,提出RedEvoAgent,通过提炼攻击轨迹、自适应进化技能和验证机制来自动红队测试,优于现有方法。