๐ค AI Summary
Evaluating the robustness of video classification models faces significant challenges due to high computational complexity and cost introduced by the temporal dimension, particularly in achieving efficient misclassification while minimizing perturbation magnitude. This paper proposes a spatiotemporal-decoupled multi-agent reinforcement learning (MARL) attack framework: two cooperative agentsโone spatial, one temporalโare jointly trained to identify sensitive spatiotemporal regions, thereby generating temporally coherent, visually imperceptible, fine-grained adversarial perturbations. The framework supports customizable distortion types and simultaneously enforces โโ-norm constraints and query efficiency. Experiments on HMDB-51 and UCF-101 demonstrate that our method significantly outperforms state-of-the-art approaches in both โโ distance and average query count across four mainstream video action recognition models. To the best of our knowledge, this is the first work enabling efficient, controllable, and fine-grained robustness evaluation for video classification.
๐ Abstract
Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101.