Score
Designs and implements evaluation protocols and adaptive adversarial attack pipelines that are explicitly aware of deployed defenses and that iterate against those defenses under a specified threat model and protocol. Analyzes and measures attack success and system degradation under adaptation, reproduces and extends prior analyses to compare defenses, and produces metrics and procedures for defense-aware red teaming.
Traditional manual red-teaming struggles to meet the demands of modern AI applications for efficient and scalable security evaluation. This work presents the first systematic survey of algorithmic red-teaming approaches tailored for AI systems, synthesizing key techniques—including AI-driven attack simulation, automated vulnerability discovery, and adversarial testing frameworks—through a comprehensive literature analysis. The study establishes a unified methodological framework and tool ecosystem, delineates the current scope and limitations of the field, identifies critical research gaps, and outlines promising future directions. By doing so, it provides both theoretical foundations and a practical roadmap to enhance the efficiency, adaptability, and comprehensiveness of security assessments for AI applications.
Large language models (LLMs) pose increasingly critical privacy, security, and ethical risks, necessitating proactive, systematic safety evaluation. Method: This paper introduces the first end-to-end, multi-component LLM red-teaming system architecture for actively identifying and quantitatively assessing generative AI safety vulnerabilities. It integrates prompt injection, adversarial example generation, automated testing frameworks, and standardized safety benchmarks (e.g., HarmBench) into a unified pipeline covering attack generation, success-rate measurement, and multidimensional evaluation—namely, effectiveness, generalizability, and reproducibility. Contributions/Results: Key innovations include a systematic safety-enhancement framework, an open-source toolchain integration strategy, a reusable red-teaming practice guide, and a tool-selection matrix. Empirical evaluation demonstrates that the framework significantly improves developers’ efficiency in risk identification and enables reliable GenAI deployment in high-assurance settings.
This work proposes a general-purpose red-teaming framework that overcomes the limitations of existing automated approaches, which are often confined to specific security scenarios and rely on evaluators known during training, thereby lacking generalization to novel adversarial targets. By end-to-end fine-tuning compact language models such as Qwen3-8B and integrating multi-objective adversarial example generation with adaptive optimization strategies, the method generates effective attacks against arbitrary red-teaming tasks without requiring predefined evaluators. Experimental results demonstrate significant improvements in attack generation performance both within and across domains. To the best of our knowledge, this is the first approach to achieve evaluator-agnostic, generalizable red-teaming automation, effectively transcending the constraints of conventional methods in terms of task scope and adaptability.
This work addresses the limitations of existing automated red-teaming approaches, which are constrained by manually designed workflows and struggle to efficiently explore the system design space. The authors propose a novel framework that, for the first time, formulates red-teaming as an agent-based system design problem. By integrating in-context learning with large language models and an evolutionary selection mechanism, the framework enables end-to-end autonomous evolution of red-team system architectures without human intervention. This approach overcomes the conventional paradigm of optimizing attack strategies within fixed structures. Empirical results demonstrate remarkable performance: achieving 96% and 98% attack success rates on Llama-2-7B and Llama-3-8B, respectively, and 100% success on both GPT-3.5-Turbo and GPT-4o-mini, significantly outperforming prior methods and exhibiting strong cross-model transferability.
This work addresses the limited robustness of existing AI-driven SOAR systems against adaptive adversaries and the absence of effective evaluation methodologies. To this end, the authors propose an autonomous red teaming framework that integrates the strategic planning capabilities of large language models (LLMs) with the tactical execution strengths of reinforcement learning (RL) through a hierarchical hybrid architecture. A kill-chain-aligned reward mechanism is designed to generate multi-stage, adaptive attacks within a high-fidelity enterprise network simulation environment. Experimental results demonstrate that pure LLM-based agents struggle to sustain prolonged attacks, and specialized security models achieve only limited breaches. In contrast, the proposed LLM-RL hybrid approach significantly enhances attack simulation fidelity and effectively evaluates the defensive resilience of SOAR systems.
Traditional AI red-teaming relies on manual, task-specific procedures that are time-consuming and difficult to reuse, thereby limiting the efficiency of security evaluations. This work proposes a new red-teaming paradigm for the agent era: a natural language–driven AI red teaming agent built on the Dreadnode SDK that automatically orchestrates and executes end-to-end testing workflows encompassing attacks, transformations, and scoring. The framework unifies security assessment for both traditional machine learning and generative AI systems, supporting multi-agent, multilingual, and multimodal targets. It enables access to over 45 attack strategies, 450 transformations, and 130 scorers without requiring manual coding. In a case study with Meta’s Llama Scout, natural language instructions alone achieved an 85% attack success rate (severity 1.0), reducing testing cycles from weeks to hours.
This paper addresses the challenge of identifying and mitigating security threats across the full lifecycle of large language models (LLMs). To this end, it proposes the first structured threat model explicitly aligned with the LLM development-to-deployment pipeline. Methodologically, it integrates threat modeling, a systematic mapping of knowledge (SoK), inductive analysis of attack patterns, and red-teaming practice to construct a comprehensive, stage-specific attack taxonomy—characterizing key attacker motivations, entry points, and corresponding defensive countermeasures. Its primary contribution is the first LLM-specific, phase-aware attack classification framework, which underpins a reusable, operationally grounded red-teaming methodology. This framework significantly enhances the systematicity, practicality, and industrial applicability of LLM security assessments, providing both theoretical foundations and actionable guidance for robust LLM security hardening.
Addressing the challenges of safety alignment for large language models (LLMs) in high-stakes applications and the poor scalability of manual red-teaming, this paper introduces the first meta-prompt-driven automated red-teaming framework. The framework integrates multimodal anomaly detection with structured threat modeling to enable closed-loop evaluation across six standardized threat categories—including Reward Hacking and Deceptive Alignment. It pioneers a meta-prompt-guided adversarial prompt synthesis mechanism and a novel multimodal vulnerability co-detection paradigm, uncovering 12 previously undocumented attack patterns. Evaluated on GPT-OSS-20B, the framework identifies 47 vulnerabilities—including 21 high-severity ones—with a 3.9× higher detection rate than human experts and 89% accuracy. This significantly improves reproducibility, coverage, and interpretability in AI safety testing.
This study addresses critical security vulnerabilities prevalent in current offensive AI agent systems, which lack systematic evaluation frameworks. The work proposes the first comprehensive attack-chain model encompassing large language model (LLM) manipulation, lateral movement, persistence, defense evasion, and sandbox escape, thereby uncovering common architectural flaws. By integrating red-teaming methodologies, LLM security analysis, and container escape detection, the authors reproduce and validate multiple high-severity vulnerabilities—including API key exfiltration and host machine compromise. Building on these findings, they formulate a set of architecture-level, broadly applicable security design principles that effectively mitigate the identified attack vectors and substantially enhance the overall system resilience.
This study addresses the challenge of malicious evolution in agent skills under dual defense mechanisms comprising pre-execution scanning and runtime protection. To this end, it proposes a two-stage feedback loop framework that introduces a novel cross-stage closed-loop evolution mechanism. Specifically, defensive feedback from both stages is transformed into adaptive red-teaming learning signals. By integrating scanner-guided optimization, task-conditioned verification, and large language model-based code generation, the framework enables the automated evolution of complete malicious skill packages. Experimental results demonstrate that the proposed method achieves an average attack success rate of 45.28%, outperforming baselines by approximately five percentage points. Furthermore, it effectively evades security scanners while preserving benign task performance.
This study investigates the use of large language model agents (LMAs) to automate advanced persistent threat behaviors—such as lateral movement—in red teaming exercises, while addressing reliability and governance challenges inherent in real-world adversarial environments. Grounded in the MITRE ATT&CK framework, we evaluate LMA performance across three operational modes—fully autonomous execution, self-guided planning, and expert-defined playbooks—within a controlled adversarial simulation environment. Agents interact with instrumented network proxies, observe execution traces, and iteratively adapt based on feedback. We introduce, for the first time in red teaming, an LLM-as-a-Judge paradigm to deterministically validate attack chains and systematically delineate LMA capability boundaries. Results show that expert-defined playbooks achieve the highest completion rates, yet all modes frequently fail, primarily due to fragile command invocation, environmental instability, and errors in credential and state management.
This work addresses the lack of effective evaluation against adaptive, defense-aware attacks in existing defenses against indirect prompt injection for large language model agents. It is the first to unify out-of-band defense mechanisms within classical security frameworks—specifically Biba integrity, reference monitors, and least privilege—and systematically analyzes their capabilities alongside information flow label design. The study further introduces the first adaptive attack evaluation protocol tailored to such defenses. Using Qwen2.5-7B deployed on a single H200 GPU, the authors reproduce and extend the Progent attack, conducting black-box evaluations on the AgentDojo benchmark. Results demonstrate that Progent’s defense reduces average attack success rates from 25.8% to 4.2%, with handcrafted adaptive attacks achieving only 2.6%, thereby confirming the substantial robustness of out-of-band defenses against dynamic adversarial strategies.