adversarial training

Design, build, and evaluate machine-learning training pipelines, models, and defenses that maintain performance under intentional manipulation, including methods that synthesize and model adversarial perturbations and strategies, and incorporate adversarial examples into optimization (adversarial training). This competence also covers devising adversarial strategy models, selecting robust losses and architectures, mitigating adversary-induced label noise, learning from sparse or ambiguous labels, and adapting models continuously to evolving threat models.

adversarialtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Adversarial Machine Learning: Attacks, Defenses, and Open Challenges

Feb 08, 2025
PK
Pranav Kumar Jha
🏛️ AI Solutions Architect

Adversarial machine learning suffers from fundamental robustness deficiencies under evasion and poisoning attacks, undermining AI reliability in safety-critical applications. Method: We propose the first unified mathematical framework formalizing diverse attack and defense classes, explicitly characterizing the inherent tension among certified robustness, scalability, and practical deployability. Our approach integrates game-theoretic modeling, optimization-theoretic analysis, formal verification, and empirical evaluation to establish a systematic, end-to-end analytical paradigm spanning the full attack–defense spectrum. Contributions: (1) We identify theoretical and practical bottlenecks in robustness guarantees under adaptive adversaries; (2) We systematically characterize and structure three open challenges—ill-defined boundaries of certified robustness, insufficient scalability to large-scale settings, and lack of reliability in real-world deployment; (3) We provide verifiable theoretical benchmarks and principled design guidelines for next-generation robust AI systems.

Address vulnerabilities in AI systemsDiscuss challenges in robust solutionsFormalize defense mechanisms rigorously

Defending against adversarial attacks using mixture of experts

Dec 23, 2025
MM
Mohammad Meymani
🏛️ University of New Brunswick

To address the insufficient robustness of machine learning models against diverse threats—including adversarial examples, data poisoning, and model extraction—this paper proposes an end-to-end adversarial training framework based on a Mixture of Experts (MoE). The framework integrates nine ResNet-18 experts with a learnable gating mechanism and, for the first time, deeply embeds adversarial training into the MoE architecture, enabling joint optimization of expert parameters and routing policies. Compared to complex monolithic models, our approach achieves superior robustness using a significantly lighter backbone: it substantially outperforms existing defenses and stronger baseline models under standard adversarial attacks. This demonstrates the dual advantage of lightweight MoE architectures—enhanced robustness without compromising computational efficiency.

Defending machine learning models against adversarial attacksEnhancing robustness using mixture-of-experts architectureImproving performance over complex state-of-the-art systems

On the Effectiveness of Adversarial Training on Malware Classifiers

Dec 24, 2024
HB
Hamid Bostani
🏛️ Radboud University | King's College London | TU Wien | University College London | Ruhr University Bochum

Adversarial training (AT) is widely adopted for robust malware detection, yet its true robustness gains against realistic evasion attacks—while preserving high clean-sample accuracy—remain poorly understood and often overestimated. Method: We propose a comprehensive evaluation framework that decouples the coupled effects of data quality, feature representation, model architecture, and optimization strategy. It integrates static and dynamic features, multi-stage PGD/FGSM attacks, and realistic attack modeling. Contribution/Results: Our analysis reveals that AT’s effectiveness critically depends on synergistic interactions among multiple factors. We identify five common evaluation pitfalls and formulate ten reproducible, interpretable best practices for robust training. Empirically, our approach achieves >95% clean-sample accuracy while delivering significant and differentiated robustness improvements against strong, realistic evasion attacks—establishing a theoretically rigorous and engineering-practical paradigm for evaluating and optimizing security-critical AI systems.

Adversarial TrainingMalware DetectionOptimization

Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning

Mar 03, 2025
KD
Kyle Domico
🏛️ University of Wisconsin-Madison

This work addresses the low query efficiency and limited success rate of black-box model evasion attacks by being the first to systematically integrate reinforcement learning (RL) into adversarial example generation. We formulate the attack process as a Markov decision process, design state and action spaces to encode image perturbations and oracle feedback, and employ the Proximal Policy Optimization (PPO) algorithm to enable continuous policy improvement and experience reuse. The proposed method supports controllable perturbations and self-evolving attack policies, significantly enhancing both query efficiency and robustness. On CIFAR-10, it achieves a 19.4% higher attack success rate and reduces the average queries per sample by 53.2% compared to baseline methods. After 5,000 training iterations, its success rate surpasses that of SquareAttack by 13.1%.

Demonstrates RL's effectiveness in generating adversarial examples efficiently.Develops RL-based adversarial attacks on machine learning models.Improves attack success rates and reduces victim model queries.

Leveraging Generalizability of Image-to-Image Translation for Enhanced Adversarial Defense

Apr 02, 2025
HZ
Haibo Zhang
🏛️ Kyushu Institute of Technology | The University of Kitakyushu | Kyushu University

To address the vulnerability of image classification models to diverse unknown adversarial attacks and the poor generalizability and high computational overhead of existing defenses, this paper proposes a lightweight preprocessing defense framework based on image-to-image translation. Our method trains only a single residual-enhanced translation model, enabling robust adversarial purification—without fine-tuning—across attack types (e.g., FGSM, PGD, CW) and target models (e.g., ResNet, VGG, ViT). The key innovation lies in incorporating a residual architecture into the translation model to enhance cross-domain generalization, eliminating the need for attack- or model-specific customization. On a multi-attack–multi-model benchmark, our approach restores classification accuracy from near 0% to an average of 72%, while incurring significantly lower inference latency and memory footprint compared to state-of-the-art defense methods.

Enhancing adversarial defense with image-to-image translationImproving model generalizability against diverse adversarial attacksReducing training overhead while maintaining defense effectiveness

Latest Papers

What's happening recently
View more

This work systematically evaluates the impact of defensive training on large language model (LLM) agents, revealing a critical trade-off between safety alignment and functional capability. Through comprehensive assessment across 97 multi-step tasks and 1,000 adversarial prompts, the study identifies a “capability–alignment paradox”: while intended to enhance security, defensive training severely degrades agents’ task execution performance and remains vulnerable to sophisticated attacks. The authors uncover three agent-specific biases—task incompetence bias, cascading amplification bias, and trigger bias—and demonstrate that defended models fail due to timeouts in 99% of tasks (versus 13% for the baseline). Large-scale adversarial evaluation, root-cause analysis, tool-use logs, and retry behavior tracking further show that most attacks readily bypass existing defenses, indicating that current approaches sacrifice practical utility without achieving meaningful security guarantees.

Autonomy Taxcapability-alignment paradoxdefense training

This work proposes a general-purpose red-teaming framework that overcomes the limitations of existing automated approaches, which are often confined to specific security scenarios and rely on evaluators known during training, thereby lacking generalization to novel adversarial targets. By end-to-end fine-tuning compact language models such as Qwen3-8B and integrating multi-objective adversarial example generation with adaptive optimization strategies, the method generates effective attacks against arbitrary red-teaming tasks without requiring predefined evaluators. Experimental results demonstrate significant improvements in attack generation performance both within and across domains. To the best of our knowledge, this is the first approach to achieve evaluator-agnostic, generalizable red-teaming automation, effectively transcending the constraints of conventional methods in terms of task scope and adaptability.

adversarial goalsautomated red teamingcontent safety

Existing machine learning defense mechanisms primarily focus on the attacks themselves and struggle to identify the attackers, thereby limiting the effectiveness of system-level mitigation strategies. This work proposes the first domain-agnostic framework that shifts the defensive perspective from the attack to the attacker by modeling adversarial behavior and leveraging probabilistic inference to infer attacker characteristics without prior knowledge. Theoretical analysis shows that while attackers cannot be uniquely identified, their attributes can be characterized probabilistically. The framework is applicable across diverse learning models and attack scenarios. Experimental results demonstrate that it not only enhances the precision of exogenous mitigation strategies but also improves the performance of endogenous defense mechanisms such as adversarial regularization.

adversarial defenseadversary identificationattacker characteristics

This study addresses the emerging threat of intentional deception by large language model (LLM) agents in multi-agent systems, proposing a systematic framework to understand and counter such behavior. The work models intentional deception as a controllable capability and introduces a text-based RPG experimental platform encompassing 36 distinct behavioral profiles. A two-stage approach is employed: first inferring the target agent’s motivations and beliefs with high accuracy (>98%), then generating strategically misleading content to induce actions contrary to its stated stance. The findings reveal that 88.5% of successful deceptions rely on strategic reframing of factual information rather than outright fabrication, and that deception efficacy is highly dependent on the target’s specific behavioral profile. Moreover, existing fact-checking mechanisms demonstrate limited effectiveness against this form of strategic deception.

adversarial manipulationbehavioral profilesintentional deception

Existing black-box attack methods struggle to effectively evaluate the robustness of multi-component NLP systems under stringent constraints—specifically, binary feedback only, no gradient access, and a query budget of ten or fewer. This work proposes a dual-agent adversarial rewriting framework: an attack agent generates semantics-preserving rewrites, while a prompt optimization agent iteratively refines the attack strategy based solely on binary feedback. The approach achieves the first effective black-box attacks under such strict conditions, revealing critical links between system architecture and vulnerability, and identifying four distinct attack patterns targeting different pipeline stages. Experiments demonstrate evasion rates of 19.95%–40.34% against four LLM-based misinformation detection systems and up to 97.02% against static retrieval systems. Furthermore, defenses informed by these attack patterns reduce evasion rates by as much as 65.18%.

adversarial robustnessarchitectural vulnerabilitiesblack-box NLP pipelines

Hot Scholars

FR

Fabio Roli

Professor, University of Genova and Cagliari, Italy
Pattern recognitionmachine learningcomputer visioncomputer security
ZZ

Zhengyu Zhao

Xi'an Jiaotong University, China
Adversarial Machine LearningComputer Vision
XM

Xingjun Ma

Fudan University
Trustworthy AIMultimodal AIGenerative AIEmbodied AI
LD

Luca Demetrio

Assistant Professor at Università degli Studi di Genova
Computer SecurityAdversarial Machine Learning