Score
Designs, builds, and analyzes algorithmic methods and concrete adversarial inputs—perturbations, crafted prompts, and constructed instances—intended to cause model failure or to probe worst‑case model behavior. This work includes optimizing and generating attacks under constraints (e.g., quantization, reduced precision, hardware transfer), producing transferable or group‑conditional adversarial sets across modalities, evaluating model robustness under attack, and creating adversarial samples for training or analysis.
This work addresses the dual challenges of poor robustness in post-hoc interpretability methods (e.g., Grad-CAM, Integrated Gradients) and the non-traceability of conventional adversarial attacks. We propose a novel algebraic adversarial attack paradigm grounded in geometric deep learning. Methodologically, we are the first to model the symmetry group of neural networks as a Lie group, enabling algebraic characterization of invariance and fragility in explanation models—thus yielding analytically tractable and provably traceable adversarial examples. Unlike optimization-driven black-box attacks, our approach provides a mathematically verifiable generation mechanism. Experiments on CIFAR-10, an ImageNet subset, and real-world medical imaging datasets demonstrate that the proposed attack significantly degrades both faithfulness and stability of mainstream explanation methods, validating its theoretical soundness and practical applicability.
Adversarial machine learning suffers from fundamental robustness deficiencies under evasion and poisoning attacks, undermining AI reliability in safety-critical applications. Method: We propose the first unified mathematical framework formalizing diverse attack and defense classes, explicitly characterizing the inherent tension among certified robustness, scalability, and practical deployability. Our approach integrates game-theoretic modeling, optimization-theoretic analysis, formal verification, and empirical evaluation to establish a systematic, end-to-end analytical paradigm spanning the full attack–defense spectrum. Contributions: (1) We identify theoretical and practical bottlenecks in robustness guarantees under adaptive adversaries; (2) We systematically characterize and structure three open challenges—ill-defined boundaries of certified robustness, insufficient scalability to large-scale settings, and lack of reliability in real-world deployment; (3) We provide verifiable theoretical benchmarks and principled design guidelines for next-generation robust AI systems.
This work investigates the theoretical mechanisms underlying cross-model transferability of adversarial examples in ensemble-based adversarial attacks. Method: We introduce the notion of “transfer error” and rigorously decompose it— for the first time—into two quantifiable components: model vulnerability and ensemble diversity. Leveraging Rademacher complexity and information-theoretic analysis, we derive a tight upper bound on transfer error and extract three actionable guidelines for error suppression. Our approach integrates theoretical analysis with multi-model collaborative optimization. Results: Extensive evaluation across 54 heterogeneous models demonstrates that the proposed framework significantly improves adversarial transfer success rates. It establishes a novel paradigm for deep learning robustness assessment, uniquely bridging theoretical rigor with practical deployability.
Adversarial examples severely compromise the robustness of deep neural networks (DNNs), yet existing defenses either rely on attack-specific priors or require architectural modifications, suffering from poor generalizability and computational overhead. This paper proposes a universal, lightweight adversarial sample detection framework that requires no model fine-tuning or additional training. Leveraging statistical distribution analysis of layer-wise activations—including gradient sensitivity modeling and anomaly detection—it enables plug-and-play detection across diverse architectures and modalities (image, video, audio). Our key contribution is the first statistically grounded, assumption-free, and interpretable detection paradigm, derived purely from intrinsic activation statistics, eliminating dependence on prior knowledge of attack types. Extensive evaluation demonstrates >95% detection accuracy across multiple datasets and attack settings, with inference overhead under 0.5% of the original model’s computational cost—significantly outperforming state-of-the-art methods.
Adversarial examples in black-box attacks often exhibit weak transferability across models. Method: This paper proposes BayAtk, a Bayesian inference-based transferable adversarial attack method. It introduces Bayesian modeling to adversarial transferability analysis for the first time, uncovering the inherent uncertainty in transferability. BayAtk designs two transferability-enhancing priors and incorporates an instance-adaptive dynamic weighting mechanism to jointly optimize perturbation direction and magnitude. Contribution/Results: Evaluated on undefended and state-of-the-art defended black-box models, BayAtk significantly outperforms existing SOTA methods, achieving an average 12.7% improvement in transfer success rate. It establishes a more robust and interpretable attack paradigm for AI security evaluation.
This work addresses robustness challenges of deep learning models in safety-critical applications, tackling three adversarial threats: (1) adversarial examples in computer vision, (2) out-of-distribution generalization (i.e., domain generalization), and (3) jailbreaking attacks against large language models (LLMs). We propose a unified robustness enhancement framework comprising: (i) a certifiably robust defense against adversarial perturbations; (ii) a cross-domain robust training paradigm grounded in out-of-distribution generalization and invariant representation learning; and (iii) an LLM jailbreaking defense integrating formal verification with prompt-attack modeling. Our approach synergistically combines adversarial training, invariance regularization, verification-driven optimization, and controllable decoding. Evaluated on medical image analysis, molecular structure recognition, and standard image classification benchmarks, it achieves state-of-the-art generalization performance. Moreover, it significantly improves jailbreaking resistance across multiple open-source LLMs, demonstrating effectiveness and scalability in multimodal and multi-task settings.
This work investigates the extent to which adversarial attacks reflect a model’s actual robustness under random noise of comparable magnitude, rather than merely characterizing worst-case scenarios. To this end, the authors propose a directional bias perturbation framework governed by a concentration parameter κ, which interpolates smoothly between isotropic noise and adversarial directions. They further introduce a novel attack strategy designed to better approximate realistic statistical noise. Through systematic evaluations on ImageNet and CIFAR-10, the study delineates the conditions under which common adversarial attacks effectively capture noise-induced failure risks, thereby offering both theoretical grounding and practical guidance for safety-oriented robustness evaluation of machine learning models.
Existing theoretical frameworks fail to distinguish whether adversarial examples exploit fragile yet predictable non-robust features in data, leading to biased robustness evaluations. This work addresses this gap by formally categorizing adversarial examples into two types: those that rely on non-robust features and those that do not. The authors propose a novel ensemble-based metric to quantify the extent to which adversarial perturbations manipulate non-robust features. By integrating adversarial attack generation with robustness analysis, the proposed framework elucidates the mechanism through which sharpness-aware minimization enhances model robustness and explains the performance discrepancy between standard and adversarial training on robust datasets. This approach offers a refined perspective for evaluating and understanding model robustness.
This work addresses the insufficient robustness evaluation of machine learning models in multilingual settings. We propose a novel adversarial attack method grounded in high-perplexity text perturbation. Methodologically, it integrates semantics-preserving word-level and sentence-level perturbations to maximize language model perplexity while maintaining textual naturalness. Notably, we introduce Bengali—the first low-resource language—into the adversarial attack framework, constructing the first Bengali adversarial dataset and validating its efficacy against cross-lingual transfer models. Experiments demonstrate that our approach significantly degrades the accuracy of state-of-the-art text classification and NLU models across multiple benchmarks (average drop of 28.6%), with low generation cost and strong transferability. Key contributions are: (1) formalizing and realizing a perplexity-driven adversarial example generation paradigm; and (2) extending adversarial attacks to resource-constrained languages, thereby advancing multilingual robustness research.
Although large language models (LLMs) exhibit remarkable capabilities, they remain vulnerable to adversarial prompt attacks that can circumvent alignment-based defenses. This work presents the first formal game-theoretic model of the strategic interaction between attackers and defenders in this context, revealing an inherent advantage for the attacker. By integrating game theory, adversarial prompt modeling, and equilibrium analysis, the study derives theoretically provable optimal strategies for both attack and defense. Experimental results across multiple mainstream LLMs and benchmark datasets demonstrate that the proposed theoretically optimal attack strategy substantially outperforms existing methods, offering a rigorous theoretical foundation and practical guidance for the secure deployment of LLMs.