Algebraic Adversarial Attacks on Explainability Models

📅 2025-03-16
📈 Citations: 0
Influential: 0
📄 PDF

career value

192K/year
🤖 AI Summary
This work addresses the dual challenges of poor robustness in post-hoc interpretability methods (e.g., Grad-CAM, Integrated Gradients) and the non-traceability of conventional adversarial attacks. We propose a novel algebraic adversarial attack paradigm grounded in geometric deep learning. Methodologically, we are the first to model the symmetry group of neural networks as a Lie group, enabling algebraic characterization of invariance and fragility in explanation models—thus yielding analytically tractable and provably traceable adversarial examples. Unlike optimization-driven black-box attacks, our approach provides a mathematically verifiable generation mechanism. Experiments on CIFAR-10, an ImageNet subset, and real-world medical imaging datasets demonstrate that the proposed attack significantly degrades both faithfulness and stability of mainstream explanation methods, validating its theoretical soundness and practical applicability.

Technology Category

Application Category

📝 Abstract
Classical adversarial attacks are phrased as a constrained optimisation problem. Despite the efficacy of a constrained optimisation approach to adversarial attacks, one cannot trace how an adversarial point was generated. In this work, we propose an algebraic approach to adversarial attacks and study the conditions under which one can generate adversarial examples for post-hoc explainability models. Phrasing neural networks in the framework of geometric deep learning, algebraic adversarial attacks are constructed through analysis of the symmetry groups of neural networks. Algebraic adversarial examples provide a mathematically tractable approach to adversarial examples. We validate our approach of algebraic adversarial examples on two well-known and one real-world dataset.
Problem

Research questions and friction points this paper is trying to address.

Develops algebraic adversarial attacks on explainability models.
Explores conditions for generating adversarial examples in explainability models.
Validates algebraic adversarial examples on multiple datasets.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Algebraic approach to adversarial attacks
Analysis of neural network symmetry groups
Mathematically tractable adversarial examples