Score
Designs and implements planning-based systems that model adversary behavior as attack graphs or connectivity graphs and synthesize concrete attack paths and exploit chains (including command-composition exploits and composed exploit sequences). Builds algorithms and tools to automatically enumerate, prioritize, and produce diverse, sound attack scenarios and sequences, and to compute metrics or guarantees for ranking and evaluating the generated paths.
Current cybersecurity situational awareness approaches suffer from incomplete attack graph modeling, low interpretability of defensive strategies, and insufficient human–machine collaborative analysis capabilities. Method: This paper proposes a security situational assessment framework integrating AI planning and causal inference. It automatically compiles network configurations and vulnerabilities into PDDL planning models; constructs an attack connectivity hypergraph to represent attack paths under incomplete information; and incorporates formal causal reasoning to enable motive-driven, counterfactual “what-if” analysis. Contribution/Results: The framework generates semantically clear, formally verifiable hardening strategies with traceable impact attribution. Experiments demonstrate significant improvements in system administrators’ ability to systematically explore, compare, and explain defense alternatives—achieving both theoretical rigor and operational practicality.
Existing security product evaluation methods struggle to model multi-step Advanced Persistent Threat (APT) attacks and lack end-to-end interpretable simulation. Method: This paper proposes Aurora—the first automated framework that formalizes attack-chain simulation as a PDDL planning problem. It leverages LLM-driven semantic parsing and knowledge distillation of threat intelligence to automatically map APT reports to structured attack models, and integrates external penetration tools to enable cross-platform, fine-grained, and traceable end-to-end simulation. Contribution/Results: Aurora—open-sourced—is validated in real-world environments against 12 MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). It achieves a 67% improvement in planning success rate and reduces average simulation time by 82%, significantly enhancing the ecological validity and trustworthiness of defensive product evaluation.
Existing MulVal tools lack the capability to model and reason about threat propagation within software supply chains (SSCs), rendering them ineffective for analyzing complex, real-world incidents such as the XZ backdoor and the 3CX dual-supply-chain attack. Method: This paper proposes a MulVal extension framework tailored for supply chain security. It introduces a novel predicate system seamlessly integrated into MulVal’s syntax, enabling—for the first time—bidirectional logical reasoning between network-level attack graphs and SSC-specific constructs, including asset dependencies, interactions, and compromise states. By incorporating facts and rules representing supply chain assets, dependency relationships, protective mechanisms, and initial configurations, the framework constructs a multi-granularity attack graph spanning both network and supply chain layers. Contribution/Results: The approach successfully reconstructs multiple real-world supply chain attack paths, significantly enhancing threat traceability, interpretability, and detection accuracy.
This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.
Existing network attack models struggle to capture cyclic dependencies, dynamic propagation, and multi-granularity causal relationships. To address this, we propose a novel probabilistic influence propagation model that unifies directed weighted attack graphs and causal graphs. The model supports cyclic structures and self-avoiding attack chain reasoning, and—uniquely—integrates probabilistic graphical models with stochastic processes into heterogeneous networked security analysis, enabling joint fine-grained inference across vulnerabilities, services, and exploitabilities. Experimental evaluation on two attack graph benchmarks and one causal graph demonstrates that our model generates quantifiable threat-path prioritization metrics and scalable structural summaries for large graphs. These capabilities significantly improve analysts’ efficiency in reasoning about complex attack chains and enhance decision-making accuracy.
This study addresses the unclear impact of symbolic predicate granularity on the effectiveness, computational cost, and fidelity of automated attack chain generation. For the first time, we empirically investigate this issue by translating 16 Atomic Red Team techniques into PDDL predicates at two levels of granularity—five and nine predicate classes—using large language models, followed by deterministic planning with Fast Downward. Our experiments reveal that in 81.3% of cases, different granularity schemes yield identical attack chains, indicating that planning effectiveness and computational cost are largely insensitive to predicate granularity. Instead, higher granularity primarily enhances the interpretability of internal plan structure without substantially affecting feasibility. This work delineates the practical boundaries of predicate granularity’s influence in modeling attack action linkages.
This work addresses the limitation of current cyber threat intelligence (CTI) reports, which lack explicit modeling of preconditions and state changes associated with attack steps, thereby hindering automated reachability analysis of multi-stage attack chains. To overcome this, the authors propose the first approach that explicitly models preconditions and postconditions during CTI extraction, representing each step as a structured attack unit comprising precondition, action, and postcondition. A large language model–driven multi-stage pipeline is employed to extract, normalize, and repair dependencies among these units, which are then compiled into Datalog rules for logical reasoning. Evaluation on 20 CTI reports containing 334 manually annotated steps demonstrates superior coverage in attack behavior extraction compared to existing systems; logical reasoning successfully reached the target in 19 reports, and backward search generated 34 valid attack paths.
This work addresses the limitations of traditional automated approaches that extract only volatile indicators—such as IP addresses, domains, and file hashes—from cyber threat intelligence (CTI), resulting in rapidly obsolete detection rules. To overcome this, the authors propose a GraphRAG-based knowledge graph-enhanced retrieval framework that, for the first time, integrates graph-structured semantic information into the automated generation of SOC hunting plans. By combining large language models with unified prompt engineering, the method extracts persistent, high-order tactical cues from CTI reports to construct robust detection logic. Experimental evaluation on nine real-world CTI reports demonstrates that, even after all underlying indicators have been rotated, the proposed approach maintains 100% detection efficacy—significantly outperforming conventional vector-retrieval RAG, which achieves only 29%—thereby substantially enhancing coverage of adversaries’ advanced tactical behaviors.