🤖 AI Summary
The vulnerability of machine learning models to adversarial attacks remains a critical security challenge. Existing deterministic defenses—such as adversarial training—overlook the inherent uncertainty in attacker behavior, while stochastic defenses often lack statistical rigor and explicit modeling assumptions. This paper introduces the first formal Bayesian framework that models adversarial uncertainty as a stochastic channel, unifying both proactive (training-time) and reactive (inference-time) defense mechanisms. The framework explicitly specifies probabilistic assumptions, providing a coherent theoretical foundation for classical methods—including adversarial training and input purification—and enabling principled co-design of proactive and reactive strategies. Experiments demonstrate that explicit modeling of adversarial uncertainty substantially improves model robustness. The framework is empirically validated across CIFAR-10, CIFAR-100, and ImageNet, confirming its effectiveness, generalizability, and scalability.
📝 Abstract
The vulnerability of machine learning models to adversarial attacks remains a critical security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-case loss. However, these deterministic approaches do not account for uncertainty in the adversary's attack. While stochastic defenses placing a probability distribution on the adversary exist, they often lack statistical rigor and fail to make explicit their underlying assumptions. To resolve these issues, we introduce a formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating all probabilistic assumptions. This yields two robustification strategies: a proactive defense enacted during training, aligned with adversarial training, and a reactive defense enacted during operations, aligned with adversarial purification. Several previous defenses can be recovered as limiting cases of our model. We empirically validate our methodology, showcasing the benefits of explicitly modeling adversarial uncertainty.