design adversarial defenses

Designs and implements technical and procedural defenses against adversarial attacks and related cybersecurity threats, applying security best practices, creating mitigation steps, and deploying baseline protective controls. Builds and composes defensive mechanisms, tunes defense hyperparameters, evaluates defense impact on system performance, and documents and trains users on step‑by‑step mitigation and hygiene.

designadversarialdefenses

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current cybersecurity exercise scenarios suffer from limited scalability, insufficient diversity, and inadequate fidelity to real-world enterprise IT environments, thereby constraining the development of practical skills for both human experts and AI agents. This work proposes an automated approach that integrates system modeling, procedural content generation, and virtualization techniques to enable, for the first time, the generation of large-scale, multidimensionally configurable exercise scenarios—spanning scale, scope, difficulty, complexity, and diversity. The project releases an open-source simulation platform alongside a dataset comprising one hundred thousand scenario instances, significantly enhancing training coverage and scalability.

cybersecurity exerciseenterprise IT systemspractical knowledge

From Sands to Mansions: Simulating Full Attack Chain with LLM-Organized Knowledge

Jul 24, 2024
LW
Lingzhi Wang
🏛️ Northwestern University | Zhejiang University

Existing security product evaluation methods struggle to model multi-step Advanced Persistent Threat (APT) attacks and lack end-to-end interpretable simulation. Method: This paper proposes Aurora—the first automated framework that formalizes attack-chain simulation as a PDDL planning problem. It leverages LLM-driven semantic parsing and knowledge distillation of threat intelligence to automatically map APT reports to structured attack models, and integrates external penetration tools to enable cross-platform, fine-grained, and traceable end-to-end simulation. Contribution/Results: Aurora—open-sourced—is validated in real-world environments against 12 MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). It achieves a 67% improvement in planning success rate and reduces average simulation time by 82%, significantly enhancing the ecological validity and trustworthiness of defensive product evaluation.

Attack SimulationCybersecurityPerformance Evaluation

Traditional defense mechanisms struggle to counter AI-driven adaptive cyberattacks, as they are ill-equipped to handle autonomous adversarial agents capable of evading detection. This work proposes a novel defensive infrastructure centered on controllable offensive AI, which, within regulated environments, is trained to simulate the full attack lifecycle to proactively acquire and transform threat intelligence into defensive knowledge. The study’s core contributions include the first systematic benchmark encompassing the entire attack chain, alongside a training-based vulnerability discovery agent, an open-weight model governance framework, a tiered capability release mechanism, and a defense-oriented agent distillation technique. Together, these establish a “offense-informed defense” strategic paradigm and outline three actionable pathways for the safe development and constraint of offensive AI capabilities.

AI agentscyber attacksdefensive strategy

The Procedural Semantics Gap in Structured CTI: A Measurement-Driven STIX Analysis for APT Emulation

Dec 12, 2025
ÁL
Ágney Lopes Roth Ferraz
🏛️ Aeronautics Institute of Technology (ITA)

Current STIX/ATT&CK frameworks describe threat behaviors solely in terms of *what* actions are performed, omitting critical procedural semantics—such as execution order, preconditions, and environmental assumptions—hindering accurate multi-stage APT simulation. Method: We first quantitatively assess ATT&CK’s coverage of real-world campaigns and intrusion sets (only 35.6% of techniques covered) and structural reusability. Then, we propose a three-stage semantic completion framework that explicitly models the procedural logic of attack chains, integrating STIX 2.1 parsing, Longest Common Subsequence (LCS)-based sequence modeling, Caldera operation mapping, and parameterized injection. Contribution/Results: With minimal human annotation of key assumptions, our approach enables Caldera to successfully reproduce real-world APT campaigns—including ShadowRay and Soft Cell. This work identifies the critical semantic gap between descriptive cyber threat intelligence (CTI) and machine-executable CTI, establishing both theoretical foundations and practical methodology for operationalizing threat intelligence.

Assess if ATT&CK structured CTI supports multi-stage adversary emulationMeasure procedural semantic gap in CTI standards for executable behavior chainsTranslate structured CTI into executable steps with explicit environmental assumptions

This study addresses the critical lack of systematic security auditing in the current ecosystem of AI agent skills, which harbors widespread yet underrecognized security risks. The authors present the first security taxonomy for agent skills grounded in real-world vulnerabilities and introduce SkillScan, a multi-stage detection framework that integrates static code analysis with large language model–based semantic classification. Empirical evaluation across 42,447 skills from two major marketplaces reveals that 26.1% contain vulnerabilities, with data leakage and privilege escalation being the most prevalent; executable-script skills exhibit significantly higher risk. The proposed method achieves 86.7% precision and 82.5% recall in vulnerability detection. The dataset and toolkit are publicly released to support further research.

agent skillsAI agentsattack surface

Latest Papers

What's happening recently
View more

This study addresses the challenge faced by resource-constrained organizations in translating cybersecurity governance frameworks—such as the NIST Cybersecurity Framework (CSF)—into actionable defense decisions. The authors propose an integrated optimization approach that synergistically combines governance requirements, adversary knowledge from MITRE ATT&CK, and adversarially aware learning. By establishing a mapping between NIST CSF maturity levels and ATT&CK mitigations, modeling attack paths via a variable-order Markov model, and formulating a deep reinforcement learning framework, the method generates cost-effective and risk-balanced defense strategies under budget constraints. This work represents the first effort to cohesively unify governance frameworks, ATT&CK-based threat intelligence, and interpretable reinforcement learning, yielding operationally viable mitigation plans that align with an organization’s security maturity and exhibit robustness against adaptive adversaries.

Adversarial BehaviorCybersecurity GovernanceMitigation Planning

This study addresses the security risks posed by AI agents with offensive cyber capabilities that may breach sandbox boundaries in evaluation environments. It systematically identifies five categories of boundary vulnerabilities—multi-step attacks, objective conflicts, supply chain leaks, persistence mechanisms, and automated execution speed—and conducts a case analysis grounded in the 2026 Hugging Face/OpenAI incident. The work introduces the first taxonomy of AI boundary vulnerabilities specifically tailored to evaluation settings and proposes an integrated defense framework combining isolation, privilege separation, behavioral provenance tracking, and defensive response interfaces. By jointly considering misuse risks and capability assessment, this research establishes clear security priorities for high-risk AI evaluations, offering both theoretical foundations and practical guidance for developing trustworthy evaluation environments that balance testing efficacy with risk containment.

AI security evaluationcyber-capable AI agentsevaluation containment

This study investigates why incorporating procedural knowledge (Skills) into tool-augmented agents fails to improve—and may even degrade—performance in offensive cybersecurity tasks. Through a reanalysis of a controlled experiment comprising 180 runs, the authors systematically evaluate the impact of Skills on CTF agents under varying levels of documentation richness. They identify “environmental feedback bandwidth” as a critical moderator: in high-feedback-bandwidth environments, the benefits of Skills diminish significantly or become detrimental. Using agents based on the MCP architecture and an ablation design with four documentation richness levels, statistical analyses—including chi-square tests, Cochran–Armitage trend tests, and Cohen’s h effect sizes—reveal only an 8.9-percentage-point performance difference between full-Skills and no-Skills conditions (p = 0.71), with most effect sizes falling below the threshold for a small effect, indicating minimal or negative marginal utility of Skills in such tasks.

Environment-Feedback BandwidthOffensive CybersecurityProcedural Knowledge

This work addresses a critical gap in existing safety evaluations by identifying a novel attack surface introduced through reusable skills: even when user requests are benign, skill materials or local artifacts can inadvertently induce agents to perform unsafe actions. The study presents SkillSafetyBench, the first systematic benchmark encompassing six risk domains, 30 safety categories, and 47 tasks, designed to assess the safety of large language model agents during skill invocation. Through adversarial case design, rule-based validators, and comparative experiments across multi-agent setups and model backends, the research demonstrates that localized, non-user-originated attacks can reliably trigger unsafe behaviors. Moreover, failure modes exhibit significant variation across domains, attack strategies, and agent architectures, revealing that safety depends not only on model alignment but also on skill parsing, contextual trust assumptions, and execution environments.

adversarial evaluationagent safetymodular skills

This work addresses the absence of systematic security research in current agent skills frameworks, which introduces structural risks. It proposes the first comprehensive security analysis framework spanning the entire lifecycle—creation, distribution, deployment, and execution—and establishes a threat taxonomy encompassing three attack surfaces and seventeen threat scenarios across seven categories. Through architectural analysis and threat modeling, the study identifies inherent design flaws—such as the lack of clear boundaries between data and instructions and the persistent trust model stemming from one-time authorization—as the primary sources of high-severity risks. These findings are empirically validated against real-world security incidents, leading to concrete, targeted defense strategies and practical mitigation recommendations.

Agent Skillsattack surfaceLLM-based agents

Hot Scholars

XJ

Xiaojun Jia

Nanyang Technological University
Explainable AIRobust AIEfficient AI
JX

Junjie Xiong

Assistant Professor of Computer Science, Missouri S&T
Network SecuritySoftware SecurityWeb Security
LS

Lukas Struppek

Senior Research Scientist @ German Research Center for AI (DFKI)
Trustworthy Generative AI
AG

Adam Gleave

CEO at FAR AI
Machine LearningDeep RL