Score
Designs, builds, and analyzes adversarial assessment artifacts and activities that probe, exploit, and evaluate system defenses, including test plans, threat-emulation scenarios, attack playbooks, evaluation frameworks, and metrics. Develops and operationalizes red-team methodologies, automated tools, coordination processes, exercises, and assessments to reveal vulnerabilities and measure resilience.
Addressing the challenges of safety alignment for large language models (LLMs) in high-stakes applications and the poor scalability of manual red-teaming, this paper introduces the first meta-prompt-driven automated red-teaming framework. The framework integrates multimodal anomaly detection with structured threat modeling to enable closed-loop evaluation across six standardized threat categories—including Reward Hacking and Deceptive Alignment. It pioneers a meta-prompt-guided adversarial prompt synthesis mechanism and a novel multimodal vulnerability co-detection paradigm, uncovering 12 previously undocumented attack patterns. Evaluated on GPT-OSS-20B, the framework identifies 47 vulnerabilities—including 21 high-severity ones—with a 3.9× higher detection rate than human experts and 89% accuracy. This significantly improves reproducibility, coverage, and interpretability in AI safety testing.
This study addresses critical security vulnerabilities prevalent in current offensive AI agent systems, which lack systematic evaluation frameworks. The work proposes the first comprehensive attack-chain model encompassing large language model (LLM) manipulation, lateral movement, persistence, defense evasion, and sandbox escape, thereby uncovering common architectural flaws. By integrating red-teaming methodologies, LLM security analysis, and container escape detection, the authors reproduce and validate multiple high-severity vulnerabilities—including API key exfiltration and host machine compromise. Building on these findings, they formulate a set of architecture-level, broadly applicable security design principles that effectively mitigate the identified attack vectors and substantially enhance the overall system resilience.
To address the fragmentation, dependency conflicts, complex deployment, and high expertise barriers associated with existing AI red-teaming tools, this paper introduces ART (AI Red-Teaming Toolkit), the first standardized, containerized platform for AI security assessment. Built on Docker, ART ensures version consistency and environment isolation, integrates 14 widely adopted open-source AI security testing tools, and employs a modular architecture to support extensibility. Inspired by the Kali Linux paradigm, it provides a unified command-line interface and cross-platform deployment capabilities (local and cloud). Its core contribution lies in the first systematic integration of heterogeneous red-teaming tools into a single, reproducible, containerized framework—significantly lowering the barrier to AI model security evaluation (including both large language models and traditional ML models), improving vulnerability detection efficiency, and enhancing experimental reproducibility.
Traditional AI red-teaming relies on manual, task-specific procedures that are time-consuming and difficult to reuse, thereby limiting the efficiency of security evaluations. This work proposes a new red-teaming paradigm for the agent era: a natural language–driven AI red teaming agent built on the Dreadnode SDK that automatically orchestrates and executes end-to-end testing workflows encompassing attacks, transformations, and scoring. The framework unifies security assessment for both traditional machine learning and generative AI systems, supporting multi-agent, multilingual, and multimodal targets. It enables access to over 45 attack strategies, 450 transformations, and 130 scorers without requiring manual coding. In a case study with Meta’s Llama Scout, natural language instructions alone achieved an 85% attack success rate (severity 1.0), reducing testing cycles from weeks to hours.
This paper addresses the challenge of identifying and mitigating security threats across the full lifecycle of large language models (LLMs). To this end, it proposes the first structured threat model explicitly aligned with the LLM development-to-deployment pipeline. Methodologically, it integrates threat modeling, a systematic mapping of knowledge (SoK), inductive analysis of attack patterns, and red-teaming practice to construct a comprehensive, stage-specific attack taxonomy—characterizing key attacker motivations, entry points, and corresponding defensive countermeasures. Its primary contribution is the first LLM-specific, phase-aware attack classification framework, which underpins a reusable, operationally grounded red-teaming methodology. This framework significantly enhances the systematicity, practicality, and industrial applicability of LLM security assessments, providing both theoretical foundations and actionable guidance for robust LLM security hardening.
Traditional manual red-teaming struggles to meet the demands of modern AI applications for efficient and scalable security evaluation. This work presents the first systematic survey of algorithmic red-teaming approaches tailored for AI systems, synthesizing key techniques—including AI-driven attack simulation, automated vulnerability discovery, and adversarial testing frameworks—through a comprehensive literature analysis. The study establishes a unified methodological framework and tool ecosystem, delineates the current scope and limitations of the field, identifies critical research gaps, and outlines promising future directions. By doing so, it provides both theoretical foundations and a practical roadmap to enhance the efficiency, adaptability, and comprehensiveness of security assessments for AI applications.
This work proposes a general-purpose red-teaming framework that overcomes the limitations of existing automated approaches, which are often confined to specific security scenarios and rely on evaluators known during training, thereby lacking generalization to novel adversarial targets. By end-to-end fine-tuning compact language models such as Qwen3-8B and integrating multi-objective adversarial example generation with adaptive optimization strategies, the method generates effective attacks against arbitrary red-teaming tasks without requiring predefined evaluators. Experimental results demonstrate significant improvements in attack generation performance both within and across domains. To the best of our knowledge, this is the first approach to achieve evaluator-agnostic, generalizable red-teaming automation, effectively transcending the constraints of conventional methods in terms of task scope and adaptability.
Current red-teaming approaches for large language models (LLMs) rely heavily on manual efforts or static datasets, resulting in low efficiency and limited capacity to uncover deep-seated security vulnerabilities. This work proposes the first automatic and adaptive red-teaming framework based on Generative Flow Networks (GFlowNets), which leverages an attacker LLM to dynamically generate highly creative adversarial inputs. The framework autonomously identifies vulnerabilities in target models and quantifies their robustness without human intervention. By introducing GFlowNets into LLM red-teaming for the first time, the method outperforms existing benchmarks in English attack generation and pioneers support for automatic adversarial input generation in low-resource languages such as Turkish, substantially enhancing test coverage and evaluation efficiency.
Existing agent monitoring evaluation methods struggle to detect highly stealthy and diverse attacks, leading to an overestimation of their defensive capabilities. This work proposes a semi-automated red-teaming framework that systematically constructs diverse attack trajectories through a three-stage pipeline: strategy generation, execution, and trajectory optimization. The approach innovatively introduces an attack taxonomy to mitigate mode collapse, decomposes the attack construction process to bridge the gap between conception and execution, and leverages large language models to enable scalable, semi-automated testing. Using this framework, the authors develop MonitoringBench—a benchmark comprising 2,644 attack trajectories within BashArena—which reduces the detection rate of state-of-the-art monitors from 94.9% to 60.3%, exposing critical weaknesses in their defenses against persuasive attacks and in the calibration of risk scoring mechanisms.
This work addresses key limitations of current large language models in automated vulnerability discovery and exploitation—namely, limited interactivity, weak execution capabilities, and poor reusability of prior experience. To overcome these challenges, the authors propose a security-aware multi-agent framework that emulates real-world red team workflows by decomposing vulnerability analysis into coordinated discovery and exploitation phases. The framework establishes a closed-loop process driven by planning, execution, verification, and feedback-based iterative refinement. Innovatively integrating execution feedback, structured agent interaction, and a long-term memory mechanism, it synergistically combines domain-specific security knowledge with code-aware analysis to enable experience reuse and continuous improvement. Evaluated across multiple security benchmarks, the approach significantly outperforms strong baselines, achieving an exploitation success rate exceeding 60% and an absolute improvement of over 10% in detection accuracy.