Score
Designs, implements, and analyzes experiments, simulations, and analytical frameworks that model, craft, characterize, detect, and rank adversarial attacks and defender responses against machine learning systems. Builds adversary simulations and threat models, attack-generation and detection tools, evaluation metrics (e.g., query efficiency or gradient dependence), attacker–defender/game-theoretic analyses, and robustness-testing protocols to quantify system performance under evasion, adaptive/continual attacks, and other adversarial conditions.
Adversarial machine learning suffers from fundamental robustness deficiencies under evasion and poisoning attacks, undermining AI reliability in safety-critical applications. Method: We propose the first unified mathematical framework formalizing diverse attack and defense classes, explicitly characterizing the inherent tension among certified robustness, scalability, and practical deployability. Our approach integrates game-theoretic modeling, optimization-theoretic analysis, formal verification, and empirical evaluation to establish a systematic, end-to-end analytical paradigm spanning the full attack–defense spectrum. Contributions: (1) We identify theoretical and practical bottlenecks in robustness guarantees under adaptive adversaries; (2) We systematically characterize and structure three open challenges—ill-defined boundaries of certified robustness, insufficient scalability to large-scale settings, and lack of reliability in real-world deployment; (3) We provide verifiable theoretical benchmarks and principled design guidelines for next-generation robust AI systems.
To address the safety verification challenge for deep reinforcement learning (DRL) decision-support systems prior to deployment, this paper proposes the first explainable and intervenable adversarial analysis framework tailored for the pre-deployment phase. Methodologically, it integrates temporal sensitivity modeling with joint observation-dimension ranking and leverages a customized strategic simulation environment—CyberStrike—to generate precise temporal perturbations, enabling behavioral pattern identification and vulnerability localization. Key contributions include: (1) establishing a novel paradigm for DRL policy vulnerability assessment; (2) introducing a joint observation-temporal sensitivity analysis method; and (3) empirically demonstrating cross-algorithm and cross-architecture attack transferability. Experiments reveal that mainstream DRL policies exhibit high sensitivity to minute perturbations at critical decision steps, exposing widespread robustness deficiencies—providing actionable insights for DRL system hardening.
Existing machine learning defense mechanisms primarily focus on the attacks themselves and struggle to identify the attackers, thereby limiting the effectiveness of system-level mitigation strategies. This work proposes the first domain-agnostic framework that shifts the defensive perspective from the attack to the attacker by modeling adversarial behavior and leveraging probabilistic inference to infer attacker characteristics without prior knowledge. Theoretical analysis shows that while attackers cannot be uniquely identified, their attributes can be characterized probabilistically. The framework is applicable across diverse learning models and attack scenarios. Experimental results demonstrate that it not only enhances the precision of exogenous mitigation strategies but also improves the performance of endogenous defense mechanisms such as adversarial regularization.
Existing security product evaluation methods struggle to model multi-step Advanced Persistent Threat (APT) attacks and lack end-to-end interpretable simulation. Method: This paper proposes Aurora—the first automated framework that formalizes attack-chain simulation as a PDDL planning problem. It leverages LLM-driven semantic parsing and knowledge distillation of threat intelligence to automatically map APT reports to structured attack models, and integrates external penetration tools to enable cross-platform, fine-grained, and traceable end-to-end simulation. Contribution/Results: Aurora—open-sourced—is validated in real-world environments against 12 MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). It achieves a 67% improvement in planning success rate and reduces average simulation time by 82%, significantly enhancing the ecological validity and trustworthiness of defensive product evaluation.
Machine learning–based intrusion detection systems (ML-IDS) exhibit insufficient robustness against black-box adversarial attacks, where attackers manipulate inputs without access to model internals. Method: We propose a lightweight, proactive, and adaptive feature poisoning defense that operates without knowledge of model parameters or training data. Leveraging behavior-aware, context-adaptive network traffic profiling—integrated with change-point detection and dynamic scaling perturbations—it injects imperceptible, targeted perturbations into attacker-critical features. This disrupts the adversary’s attack loop, whether based on binary outputs or behavioral feedback. Contribution/Results: The approach is attack-agnostic, deployment-transparent, and preserves original IDS accuracy. Extensive experiments demonstrate significant reductions in attack success rates across realistic black-box scenarios—including silent probing, transfer attacks, and decision-boundary attacks—while maintaining high generalizability and strong robustness.
This study addresses the critical gap between theory and practice in AI-driven cyberattack prediction, focusing on outdated datasets, limited attack coverage, insufficient model interpretability, weak adversarial robustness, and privacy-ethical risks. Through a systematic review of over 150 benchmark datasets and more than 200 studies, the work introduces a novel multidimensional gap assessment framework based on detection impact, implementation cost, and remediation time to prioritize these challenges. The analysis identifies dataset obsolescence and adversarial robustness as the highest-priority issues, while highlighting interpretability as a cost-effective entry point in resource-constrained settings. Furthermore, the study proposes a tripartite classification of dataset quality—production-ready, research-only, and unusable—alongside a corresponding deployment roadmap, significantly enhancing the practical feasibility and robustness of AI-based cybersecurity systems.
This work addresses the challenge that machine learning systems often face concurrent threats to robustness, privacy, and fairness, while existing defenses typically target only a single risk and lack systematic evaluation under multi-defense co-deployment. To bridge this gap, the authors propose a modular framework that encapsulates 35 state-of-the-art defense techniques as containerized components, integrated within a unified, reproducible platform featuring an automated experimentation engine and a multidimensional evaluation suite. This platform enables flexible composition and joint assessment of defenses across the machine learning lifecycle. The study provides the first systematic analysis of reproducibility discrepancies and integration challenges among different defense families when deployed in combination, offering foundational support for building reliable machine learning systems that simultaneously satisfy multiple security and trustworthiness objectives.
This work proposes a lightweight reinforcement learning–based adversarial agent to address the vulnerability of machine learning–driven network intrusion detection systems (NIDS) to evasion attacks. Unlike conventional approaches that incur high computational overhead and deployment complexity, the proposed method employs offline training to learn perturbation strategies that effectively bypass NIDS without requiring online optimization. To the best of our knowledge, this is the first application of lightweight reinforcement learning to adversarial attacks against NIDS, supporting white-box, gray-box, and black-box threat models. Experimental results demonstrate a maximum attack success rate of 48.9%, with each perturbation generated in just 5.72 milliseconds and occupying only 0.52 MB of memory, thereby significantly enhancing the practicality and deployability of adversarial attacks in real-world scenarios.
This work addresses the limited robustness of existing machine learning–based network intrusion detection systems against adversarial threats such as gradient-based attacks and distributional shifts, as well as their inability to adaptively respond to diverse attack types. To overcome these limitations, the authors propose an attack-aware, multi-stage defense framework that uniquely integrates three complementary signals—ensemble disagreement, prediction uncertainty, and distributional anomaly—and incorporates a two-stage adaptive weight learning mechanism to enable differentiated responses to heterogeneous adversarial attacks. Experimental results demonstrate that the proposed method achieves an AUC of 94.2% on standard benchmarks, outperforming current adversarially trained ensemble models by 4.5% in accuracy and 9.0 points in F1 score. Notably, it maintains 94.4% accuracy under white-box adaptive attacks, significantly enhancing both robustness and generalization.