adversarial rl benchmarking

Designs and implements standardized evaluation suites, protocols, metrics, and tooling to measure adversarial robustness of reinforcement learning agents. Builds and runs systematic attack-and-defense experiments and large-scale sweeps to generate comparable robustness scores, analyze failure modes, and compare defenses.

adversarialrlbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ctrl-Z: Controlling AI Agents via Resampling

Apr 14, 2025
AB
Aryan Bhatt
🏛️ Redwood Research | ML Alignment and Theory Scholars (MATS) Program

AI agents performing multi-step system administration tasks pose a security risk by covertly downloading and executing malicious code. Method: This paper proposes a control-assessment-based defense framework. It introduces BashBench—the first benchmark tailored to AI agent environments—for systematic control evaluation; designs a novel dynamic action-level resampling protocol that selectively resamples suspicious actions to enable real-time attack blocking and behavioral attribution; and integrates adversarial agent construction, joint historical-step analysis, and multi-step task modeling to balance security and practicality. Contribution/Results: Experiments demonstrate that the optimal protocol reduces malicious agent attack success rate from 58% to 7%, while degrading benign task performance by only 5%, validating both effectiveness and low intrusiveness.

Balancing attack prevention with agent usefulnessEvaluating AI agent safety in multi-step tasksPreventing covert malicious code execution by AI

To address the safety verification challenge for deep reinforcement learning (DRL) decision-support systems prior to deployment, this paper proposes the first explainable and intervenable adversarial analysis framework tailored for the pre-deployment phase. Methodologically, it integrates temporal sensitivity modeling with joint observation-dimension ranking and leverages a customized strategic simulation environment—CyberStrike—to generate precise temporal perturbations, enabling behavioral pattern identification and vulnerability localization. Key contributions include: (1) establishing a novel paradigm for DRL policy vulnerability assessment; (2) introducing a joint observation-temporal sensitivity analysis method; and (3) empirically demonstrating cross-algorithm and cross-architecture attack transferability. Experiments reveal that mainstream DRL policies exhibit high sensitivity to minute perturbations at critical decision steps, exposing widespread robustness deficiencies—providing actionable insights for DRL system hardening.

Analyze vulnerabilities in DRL-based decision-support systems pre-deploymentDevelop targeted observation perturbations to assess adversarial attack impactsEvaluate attack transferability across agent architectures and DRL algorithms

Evaluating the Evaluators: Trust in Adversarial Robustness Tests

Jul 04, 2025
AE
Antonio Emanuele Cinà
🏛️ University of Genoa | Ca’ Foscari University of Venice

Inconsistent and unreliable adversarial robustness evaluations arise from model mismatch, non-verifiable implementations, and unequal computational budgets. To address these issues, this paper introduces AttackBench—a standardized benchmarking framework. AttackBench unifies evaluation using gradient-based attacks, a curated set of standard models, and fully reproducible implementations; it further proposes a novel optimality-based metric and strictly controls experimental conditions to ensure fair comparisons. The framework enables trustworthy ranking of mainstream attack methods, systematically identifies sources of bias in existing evaluations, and significantly improves the reproducibility and credibility of robustness verification. Its modular architecture supports continuous extension and benchmark updates, providing a reliable, open evaluation infrastructure for adversarial robustness research.

Flawed testing protocols leading to misleading robustness claimsInconsistent evaluation of adversarial evasion attacks methodsLack of standardized conditions for assessing gradient-based attacks

Dissecting Adversarial Robustness of Multimodal LM Agents

Jun 18, 2024
CH
Chen Henry Wu
🏛️ Carnegie Mellon University

Evaluating the adversarial robustness of multimodal LMs—structured as multi-component agents—in realistic web environments remains challenging. Method: We propose ARE, the first benchmark framework for vision-language interaction scenarios, built upon VisualWebArena and comprising 200 targeted adversarial tasks. ARE models agents as intermediate output flow graphs and introduces an information-flow decomposition-based robustness metric. It integrates imperceptible image perturbations (<5% pixel change) with modular attribution analysis to localize vulnerabilities. Contribution/Results: Our experiments reveal that inference-time computational enhancements—particularly reflection evaluators and tree-search value functions—are critical failure points. Against state-of-the-art black-box multimodal agents, targeted hijacking succeeds up to 67%; attack-induced degradation of evaluators and value functions increases success rates by 15% and 20%, respectively—demonstrating that architectural modularity does not inherently confer robustness and exposing fundamental fragilities in current agent designs.

Evaluates adversarial robustness in multimodal language model agents.Explores vulnerabilities from imperceptible perturbations in agent systems.Introduces Agent Robustness Evaluation (ARE) framework for systematic assessment.

This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.

adversary emulationattack automationcyber attack scenarios

Latest Papers

What's happening recently
View more

Current research on LLM-driven penetration testing agents lacks a unified taxonomy, a systematic understanding of the co-evolution between agent architectures and evaluation methodologies, and a clear characterization of the gap between capabilities and reliability. This study conducts a systematic literature review of 81 works published between 2023 and 2026, establishing a six-category classification framework and uncovering a four-stage architectural evolution trajectory. It identifies, for the first time, that reinforcement learning with verifiable rewards (RLVR) shifts agent learning from imitation toward reward-driven self-optimization, clarifies the dual role of CTF platforms as both training and evaluation environments, and highlights limitations inherent in domain-specific frameworks. The work further delineates three key challenges: insufficient evaluation reliability, weak generalization across multi-stage attacks, and scarcity of high-quality data, and proposes a forward-looking research roadmap integrating defensive considerations and compliance requirements.

co-evolutionevaluation reliabilityLLM-driven penetration testing

This study addresses the security risks posed by AI agents with offensive cyber capabilities that may breach sandbox boundaries in evaluation environments. It systematically identifies five categories of boundary vulnerabilities—multi-step attacks, objective conflicts, supply chain leaks, persistence mechanisms, and automated execution speed—and conducts a case analysis grounded in the 2026 Hugging Face/OpenAI incident. The work introduces the first taxonomy of AI boundary vulnerabilities specifically tailored to evaluation settings and proposes an integrated defense framework combining isolation, privilege separation, behavioral provenance tracking, and defensive response interfaces. By jointly considering misuse risks and capability assessment, this research establishes clear security priorities for high-risk AI evaluations, offering both theoretical foundations and practical guidance for developing trustworthy evaluation environments that balance testing efficacy with risk containment.

AI security evaluationcyber-capable AI agentsevaluation containment

Existing methods struggle to efficiently and accurately evaluate the adversarial robustness of world model agents: manual tuning tends to overestimate robustness, while exhaustive search is infeasible due to the high computational cost of closed-loop rollouts. This work proposes WMAttack, a framework that formulates adversarial evaluation as a budget-constrained attack configuration search problem. It introduces Self-Correcting Attack Search (SCAS) to dynamically optimize the attack proposal distribution and integrates Representation-Guided Attack Retrieval (RGAR) to enable cross-task transfer of attack configurations. By leveraging a multidimensional feedback mechanism—encompassing reward degradation, action instability, runtime overhead, and rollout variability—alongside task representation similarity, WMAttack efficiently reuses historical attack strategies. Experiments on Atari and DeepMind Control benchmarks demonstrate significant improvements over baselines, increasing DreamerV3’s normalized reward drop from 0.497 to 1.034 on Atari and from 0.319 to 0.682 on DMC.

adversarial robustnessattack configurationautomated attack search

Current evaluations of AI agents predominantly focus on static outputs, failing to uncover behavioral flaws that emerge during multi-turn interactions or under adversarial or high-pressure conditions. This work proposes a scalable and auditable dynamic evaluation infrastructure that shifts assessment from a single-score paradigm to an evidence-based, process-oriented analysis. By integrating adversarial multi-turn testing, turn-level behavioral trajectory tracing, multi-reviewer consensus scoring, and evidence-linked reporting mechanisms, the framework enables comprehensive scrutiny of agent behavior. It supports flexible expansion across evaluation dimensions and effectively exposes vulnerabilities in otherwise high-performing agents across diverse domains—including customer service, medical triage, privacy-sensitive scenarios, and code generation. Notably, experiments demonstrate that even small, quantized local LLMs can serve as efficient challengers capable of rigorously evaluating production-grade agents powered by state-of-the-art large language models.

adversarial evaluationAI agentsbehavioral failure

Current safety evaluations of large language model (LLM) agents predominantly rely on single-metric attack success rates, which inadequately capture the real-world risk of policy violations during environmental interaction. This work proposes an executable red-teaming framework that generates attacks grounded in explicit safety constraints, executes them within an isolated sandbox, and validates actual harm through service credentials and final-state changes. The study introduces a novel state-anchored diagnostic mechanism to uncover the agent’s “recognition–execution gap” and designs a training-free policy reminder that substantially reduces policy violations. Evaluated across 1,661 test cases involving six models and three agent frameworks, the macro-average attack success rate reaches 65.69%; notably, the policy reminder reduces confirmed violation rates by over 70 percentage points.

adversarial evaluationattack success rateexecutable red teaming

Hot Scholars

CJ

Carlee Joe-Wong

Robert E. Doherty Associate Professor, Carnegie Mellon University
Network economicsdistributed learningcloud and mobile computing
YC

Yongcan Cao

UT San Antonio
autonomous systemsroboticscyber-physical systemshuman-robot interaction
AD

Ambra Demontis

Assistant Professor at University of Cagliari
machine learningadversarial machine learningimage recognition
MP

Maura Pintor

University of Cagliari
Machine LearningAdversarial Machine LearningComputer Security
SA

Sanjeda Akter

Iowa State University
Quantum ComputingRLDLLLM