A unified Bayesian framework for adversarial robustness

📅 2025-10-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The vulnerability of machine learning models to adversarial attacks remains a critical security challenge. Existing deterministic defenses—such as adversarial training—overlook the inherent uncertainty in attacker behavior, while stochastic defenses often lack statistical rigor and explicit modeling assumptions. This paper introduces the first formal Bayesian framework that models adversarial uncertainty as a stochastic channel, unifying both proactive (training-time) and reactive (inference-time) defense mechanisms. The framework explicitly specifies probabilistic assumptions, providing a coherent theoretical foundation for classical methods—including adversarial training and input purification—and enabling principled co-design of proactive and reactive strategies. Experiments demonstrate that explicit modeling of adversarial uncertainty substantially improves model robustness. The framework is empirically validated across CIFAR-10, CIFAR-100, and ImageNet, confirming its effectiveness, generalizability, and scalability.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessGame Theory and Economic Paradigms: Adversarial Learning

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSecurity and Privacy: Security and privacy of machine learning and AI applicationsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
The vulnerability of machine learning models to adversarial attacks remains a critical security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-case loss. However, these deterministic approaches do not account for uncertainty in the adversary's attack. While stochastic defenses placing a probability distribution on the adversary exist, they often lack statistical rigor and fail to make explicit their underlying assumptions. To resolve these issues, we introduce a formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating all probabilistic assumptions. This yields two robustification strategies: a proactive defense enacted during training, aligned with adversarial training, and a reactive defense enacted during operations, aligned with adversarial purification. Several previous defenses can be recovered as limiting cases of our model. We empirically validate our methodology, showcasing the benefits of explicitly modeling adversarial uncertainty.
Problem

Research questions and friction points this paper is trying to address.

Modeling adversarial uncertainty through Bayesian stochastic channels
Addressing lack of statistical rigor in existing stochastic defenses
Providing proactive and reactive robustification strategies against attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian framework models adversarial uncertainty via stochastic channel
Proactive defense strategy aligned with adversarial training
Reactive defense strategy aligned with adversarial purification
🔎 Similar Papers
No similar papers found.
P
Pablo G. Arce
Institute of Mathematical Sciences, Spanish National Research Council, Madrid, Spain
P
Pablo G. Arce
Universidad Autónoma de Madrid, Escuela de Doctorado, Madrid, Spain
Roi Naveiro
Roi Naveiro
CUNEF Universidad
Probabilistic Machine LearningAdversarial Machine LearningBayesian StatisticsDecision Analysis
D
David Ríos Insua
Institute of Mathematical Sciences, Spanish National Research Council, Madrid, Spain