Score
Designs and implements technical and procedural defenses against adversarial attacks and related cybersecurity threats, applying security best practices, creating mitigation steps, and deploying baseline protective controls. Builds and composes defensive mechanisms, tunes defense hyperparameters, evaluates defense impact on system performance, and documents and trains users on step‑by‑step mitigation and hygiene.
Current cybersecurity exercise scenarios suffer from limited scalability, insufficient diversity, and inadequate fidelity to real-world enterprise IT environments, thereby constraining the development of practical skills for both human experts and AI agents. This work proposes an automated approach that integrates system modeling, procedural content generation, and virtualization techniques to enable, for the first time, the generation of large-scale, multidimensionally configurable exercise scenarios—spanning scale, scope, difficulty, complexity, and diversity. The project releases an open-source simulation platform alongside a dataset comprising one hundred thousand scenario instances, significantly enhancing training coverage and scalability.
Existing security product evaluation methods struggle to model multi-step Advanced Persistent Threat (APT) attacks and lack end-to-end interpretable simulation. Method: This paper proposes Aurora—the first automated framework that formalizes attack-chain simulation as a PDDL planning problem. It leverages LLM-driven semantic parsing and knowledge distillation of threat intelligence to automatically map APT reports to structured attack models, and integrates external penetration tools to enable cross-platform, fine-grained, and traceable end-to-end simulation. Contribution/Results: Aurora—open-sourced—is validated in real-world environments against 12 MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). It achieves a 67% improvement in planning success rate and reduces average simulation time by 82%, significantly enhancing the ecological validity and trustworthiness of defensive product evaluation.
Traditional defense mechanisms struggle to counter AI-driven adaptive cyberattacks, as they are ill-equipped to handle autonomous adversarial agents capable of evading detection. This work proposes a novel defensive infrastructure centered on controllable offensive AI, which, within regulated environments, is trained to simulate the full attack lifecycle to proactively acquire and transform threat intelligence into defensive knowledge. The study’s core contributions include the first systematic benchmark encompassing the entire attack chain, alongside a training-based vulnerability discovery agent, an open-weight model governance framework, a tiered capability release mechanism, and a defense-oriented agent distillation technique. Together, these establish a “offense-informed defense” strategic paradigm and outline three actionable pathways for the safe development and constraint of offensive AI capabilities.
Current STIX/ATT&CK frameworks describe threat behaviors solely in terms of *what* actions are performed, omitting critical procedural semantics—such as execution order, preconditions, and environmental assumptions—hindering accurate multi-stage APT simulation. Method: We first quantitatively assess ATT&CK’s coverage of real-world campaigns and intrusion sets (only 35.6% of techniques covered) and structural reusability. Then, we propose a three-stage semantic completion framework that explicitly models the procedural logic of attack chains, integrating STIX 2.1 parsing, Longest Common Subsequence (LCS)-based sequence modeling, Caldera operation mapping, and parameterized injection. Contribution/Results: With minimal human annotation of key assumptions, our approach enables Caldera to successfully reproduce real-world APT campaigns—including ShadowRay and Soft Cell. This work identifies the critical semantic gap between descriptive cyber threat intelligence (CTI) and machine-executable CTI, establishing both theoretical foundations and practical methodology for operationalizing threat intelligence.
This study addresses the critical lack of systematic security auditing in the current ecosystem of AI agent skills, which harbors widespread yet underrecognized security risks. The authors present the first security taxonomy for agent skills grounded in real-world vulnerabilities and introduce SkillScan, a multi-stage detection framework that integrates static code analysis with large language model–based semantic classification. Empirical evaluation across 42,447 skills from two major marketplaces reveals that 26.1% contain vulnerabilities, with data leakage and privilege escalation being the most prevalent; executable-script skills exhibit significantly higher risk. The proposed method achieves 86.7% precision and 82.5% recall in vulnerability detection. The dataset and toolkit are publicly released to support further research.
This study addresses the challenge faced by resource-constrained organizations in translating cybersecurity governance frameworks—such as the NIST Cybersecurity Framework (CSF)—into actionable defense decisions. The authors propose an integrated optimization approach that synergistically combines governance requirements, adversary knowledge from MITRE ATT&CK, and adversarially aware learning. By establishing a mapping between NIST CSF maturity levels and ATT&CK mitigations, modeling attack paths via a variable-order Markov model, and formulating a deep reinforcement learning framework, the method generates cost-effective and risk-balanced defense strategies under budget constraints. This work represents the first effort to cohesively unify governance frameworks, ATT&CK-based threat intelligence, and interpretable reinforcement learning, yielding operationally viable mitigation plans that align with an organization’s security maturity and exhibit robustness against adaptive adversaries.
This study addresses the security risks posed by AI agents with offensive cyber capabilities that may breach sandbox boundaries in evaluation environments. It systematically identifies five categories of boundary vulnerabilities—multi-step attacks, objective conflicts, supply chain leaks, persistence mechanisms, and automated execution speed—and conducts a case analysis grounded in the 2026 Hugging Face/OpenAI incident. The work introduces the first taxonomy of AI boundary vulnerabilities specifically tailored to evaluation settings and proposes an integrated defense framework combining isolation, privilege separation, behavioral provenance tracking, and defensive response interfaces. By jointly considering misuse risks and capability assessment, this research establishes clear security priorities for high-risk AI evaluations, offering both theoretical foundations and practical guidance for developing trustworthy evaluation environments that balance testing efficacy with risk containment.
This study investigates why incorporating procedural knowledge (Skills) into tool-augmented agents fails to improve—and may even degrade—performance in offensive cybersecurity tasks. Through a reanalysis of a controlled experiment comprising 180 runs, the authors systematically evaluate the impact of Skills on CTF agents under varying levels of documentation richness. They identify “environmental feedback bandwidth” as a critical moderator: in high-feedback-bandwidth environments, the benefits of Skills diminish significantly or become detrimental. Using agents based on the MCP architecture and an ablation design with four documentation richness levels, statistical analyses—including chi-square tests, Cochran–Armitage trend tests, and Cohen’s h effect sizes—reveal only an 8.9-percentage-point performance difference between full-Skills and no-Skills conditions (p = 0.71), with most effect sizes falling below the threshold for a small effect, indicating minimal or negative marginal utility of Skills in such tasks.
This work addresses a critical gap in existing safety evaluations by identifying a novel attack surface introduced through reusable skills: even when user requests are benign, skill materials or local artifacts can inadvertently induce agents to perform unsafe actions. The study presents SkillSafetyBench, the first systematic benchmark encompassing six risk domains, 30 safety categories, and 47 tasks, designed to assess the safety of large language model agents during skill invocation. Through adversarial case design, rule-based validators, and comparative experiments across multi-agent setups and model backends, the research demonstrates that localized, non-user-originated attacks can reliably trigger unsafe behaviors. Moreover, failure modes exhibit significant variation across domains, attack strategies, and agent architectures, revealing that safety depends not only on model alignment but also on skill parsing, contextual trust assumptions, and execution environments.
This work addresses the absence of systematic security research in current agent skills frameworks, which introduces structural risks. It proposes the first comprehensive security analysis framework spanning the entire lifecycle—creation, distribution, deployment, and execution—and establishes a threat taxonomy encompassing three attack surfaces and seventeen threat scenarios across seven categories. Through architectural analysis and threat modeling, the study identifies inherent design flaws—such as the lack of clear boundaries between data and instructions and the persistent trust model stemming from one-time authorization—as the primary sources of high-severity risks. These findings are empirically validated against real-world security incidents, leading to concrete, targeted defense strategies and practical mitigation recommendations.