ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of efficient methods for discovering adversarial example sets to evaluate the robustness of black-box learning systems. To this end, this work reformulates attack generation as an active learning level-set estimation problem and proposes an adaptive trust-region search framework. By integrating calibrated approximation with a local-global hybrid sampling architecture, it achieves a paradigm shift from single-point attacks to the discovery of comprehensive adversarial region sets. This approach facilitates continuous security auditing and recovers substantially larger adversarial regions under limited query budgets. Experimental evaluations on datasets such as MNIST demonstrate that the generated adversarial example sets are significantly more representative than those produced by baseline methods, including Natural Evolution Strategies (NES).
📝 Abstract
Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), which is a query-based framework that discovers adversarial input sets for black-box learning systems. ATLAS casts attack generation as an active learning level set estimation problem then combines calibrated approximations with a local-global sampling architecture to find regions of the input space that contain adversarial examples. Once discovered, ATLAS is designed to sample points within these adversarial regions to build adversarial sets that accurately represent the state of robustness of the target model. When applied on toy experiments, we find that ATLAS is able to recover more of the adversarial region under a limited query budget than does previous work. When applied to standard and adversarially trained MNIST, CIFAR, and ImageNet model targets, ATLAS produces better representative attacks than other query-based black-box attacks (NES, SignHunter, BayesOpt). ATLAS represents an automated red-teaming framework that can be used for both analyzing the robustness of learning-based systems under development and continuous auditing to see how the robustness of a system changes over time.
Problem

Research questions and friction points this paper is trying to address.

security evaluation
adversarial input sets
black-box systems
robustness assessment
active learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Search
Active Learning
Level Set Estimation
Black-box Attack
Trust-Region
🔎 Similar Papers
No similar papers found.