Score
Design, build, and evaluate models, maps, and tools that enumerate, discover, and visualize an information system’s attack surface and the assets, vulnerabilities, and attack paths that compose it. Create and run attack-simulation frameworks and multi‑turn scenario emulations (including MITRE ATT&CK mappings and attack-chain/path modeling), perform attack attribution and assessments, and derive or test defense algorithms and mitigations informed by those analyses.
Attack scenario descriptions in cybersecurity automation lack formal semantic foundations, hindering systematic analysis and automation. Method: This paper proposes an abstract, formal model based on UML class diagrams, enabling the first unified modeling of attack context and attack scenarios. The model supports structured input, automated processing, and cross-process reuse, directly facilitating two core tasks: attack analysis and automated attack script generation. Contribution/Results: Evaluated on real-world attack analysis and cybersecurity training script generation, the model demonstrates strong feasibility and effectiveness. It fills a critical gap in formal attack scenario modeling and establishes a scalable, verifiable semantic foundation for security process automation—enhancing interoperability, reproducibility, and formal reasoning in cyber defense systems.
Existing security product evaluation methods struggle to model multi-step Advanced Persistent Threat (APT) attacks and lack end-to-end interpretable simulation. Method: This paper proposes Aurora—the first automated framework that formalizes attack-chain simulation as a PDDL planning problem. It leverages LLM-driven semantic parsing and knowledge distillation of threat intelligence to automatically map APT reports to structured attack models, and integrates external penetration tools to enable cross-platform, fine-grained, and traceable end-to-end simulation. Contribution/Results: Aurora—open-sourced—is validated in real-world environments against 12 MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). It achieves a 67% improvement in planning success rate and reduces average simulation time by 82%, significantly enhancing the ecological validity and trustworthiness of defensive product evaluation.
This study addresses the challenge faced by software organizations in quantitatively evaluating the effectiveness of distinct risk management tasks against supply chain attacks. To this end, we systematically establish, for the first time, a fine-grained mapping between MITRE ATT&CK adversary tactics and techniques and tasks in the Proactive Software Supply Chain Risk Management (P-SSCRM) framework. The mapping is achieved via a rigorous four-fold independent consensus strategy to ensure reliability and validity. Beyond enabling the first formal linkage between ATT&CK and P-SSCRM, the mapping is extended to ten major security standards—including NIST SSDF and ISO/IEC 27001—facilitating cross-framework interoperability and benchmarking. The resulting attack technique–defense task matrix provides actionable, traceable mitigation pathways, significantly enhancing the precision, coverage, and coordination of supply chain risk countermeasures.
This paper addresses critical limitations of the MITRE ATT&CK framework—namely, poor real-world adaptability, weak cross-framework interoperability, and limited domain generalizability. Conducting the largest systematic review and empirical study to date (417 publications), we analyze TTP usage patterns, align ATT&CK with complementary frameworks (Cyber Kill Chain, NIST SP 800-53, STRIDE), and construct ATT&CK knowledge graph mappings. Our analysis reveals, for the first time, empirically grounded frequency distributions and effectiveness boundaries of high-utility TTPs. We propose a novel NLP- and ML-integrated ATT&CK enhancement paradigm enabling dynamic threat detection and response. Furthermore, we conduct the first large-scale empirical evaluation of ATT&CK’s applicability in critical domains—including industrial control systems and healthcare—identifying concrete adaptation bottlenecks and proposing actionable, deployable framework enhancements.
This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.
This work addresses the challenge of identifying and prioritizing multi-step attack paths in industrial control systems (ICS). The authors propose a semi-automated approach that integrates network topology and vulnerability data to construct a system model, and for the first time apply state-aware attack graph generation to a Siemens PCS7 water treatment plant blueprint. Leveraging a state-aware traversal algorithm, the method derives multi-step attack chains driven by CVEs and misconfigurations, enabling visualization of critical attack paths. Experimental results demonstrate that a single point of failure can compromise network segmentation, while remediation of key vulnerabilities effectively protects entire security zones. These findings offer actionable security insights for ICS risk mitigation.
This work proposes an end-to-end automated red-teaming framework that overcomes the limitations of existing adversary emulation approaches, which typically rely on predefined playbooks or manual intervention and struggle to automatically derive executable attack procedures from cyber threat intelligence (CTI). The proposed system uniquely integrates MITRE ATT&CK-aligned CTI report parsing, large language model–driven attack playbook generation, and an automatic execution pipeline with failure-type-aware repair mechanisms—all within a unified workflow requiring no human involvement. Implemented on the CALDERA platform and leveraging models such as Claude Sonnet 4.5, the framework generates an average of 27.3 attack capabilities per CTI report across 11 reports, achieving an 84.22% post-repair execution success rate and an F1 score of 60.50%, significantly outperforming AURORA.
This study addresses the challenge of constructing replayable systems under test (SUT) for adversary emulation using structured cyber threat intelligence (CTI), which often lacks contextual semantics. It presents the first quantitative assessment of semantic gaps in MITRE ATT&CK STIX data by integrating CAPEC and FiGHT, leveraging platform annotations, CPE/CVE linkages, and version identification to evaluate environmental representativeness. The analysis reveals that 97.6% of enterprise software objects lack version or CPE identifiers, rendering structured fields insufficient for uniquely specifying an SUT. To overcome this limitation, the work proposes a novel SUT construction paradigm that decouples corpus-supported elements from analyst assumptions, enabling the generation of multi-version-compatible and executable adversary emulation environments.
This study addresses the lack of objective validation criteria in existing threat modeling approaches, which often rely on expert judgment and are thus prone to omissions or inconsistencies. To overcome this limitation, the authors propose a quantifiable and reproducible evaluation methodology based on benchmark applications with known vulnerabilities—specifically AzureGoat and VulnBank. Using only architectural diagrams, data flow diagrams, and their textual descriptions as input, the approach evaluates the vulnerability coverage of ThreMoLIA, an LLM-assisted threat modeling system, against Microsoft Threat Modeling Tool. Experimental results demonstrate that ThreMoLIA achieves consistently higher vulnerability coverage across both benchmark applications. This work represents the first effort to employ real-world vulnerable applications as a validation benchmark for threat modeling, effectively mitigating the shortcomings inherent in traditional expert-based assessments.
开发了基于元攻击语言的MAL模拟器,用于网络攻防模拟和系统分析,并通过案例研究训练了攻防自动化代理。