Robustness Evaluation for Video Models with Reinforcement Learning

๐Ÿ“… 2025-06-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Evaluating the robustness of video classification models faces significant challenges due to high computational complexity and cost introduced by the temporal dimension, particularly in achieving efficient misclassification while minimizing perturbation magnitude. This paper proposes a spatiotemporal-decoupled multi-agent reinforcement learning (MARL) attack framework: two cooperative agentsโ€”one spatial, one temporalโ€”are jointly trained to identify sensitive spatiotemporal regions, thereby generating temporally coherent, visually imperceptible, fine-grained adversarial perturbations. The framework supports customizable distortion types and simultaneously enforces โ„“โ‚š-norm constraints and query efficiency. Experiments on HMDB-51 and UCF-101 demonstrate that our method significantly outperforms state-of-the-art approaches in both โ„“โ‚š distance and average query count across four mainstream video action recognition models. To the best of our knowledge, this is the first work enabling efficient, controllable, and fine-grained robustness evaluation for video classification.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & RobustnessMultiagent Systems: Adversarial Agents

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
๐Ÿ“ Abstract
Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101.
Problem

Research questions and friction points this paper is trying to address.

Evaluating robustness of video classification models efficiently
Minimizing perturbations to induce misclassification effectively
Enhancing attack methods with temporal and spatial coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent reinforcement learning for video robustness
Spatial-temporal coherent fine perturbations generation
Custom distortion types for relevant evaluation
๐Ÿ”Ž Similar Papers
2024-01-21International Conferences on Pattern Recognition and Artificial IntelligenceCitations: 2
Ashwin Ramesh Babu
Ashwin Ramesh Babu
Senior Research Scientist @ Hewlett Packard Labs
Computer VisionSelf-Supervised LearningAdversarial AttacksReinforcement Learningrenewable
S
Sajad Mousavi
Hewlett Packard Enterprise (Hewlett Packard Labs)
V
Vineet Gundecha
Hewlett Packard Enterprise (Hewlett Packard Labs)
S
Sahand Ghorbanpour
Hewlett Packard Enterprise (Hewlett Packard Labs)
A
Avisek Naug
Hewlett Packard Enterprise (Hewlett Packard Labs)
A
Antonio Guillen
Hewlett Packard Enterprise (Hewlett Packard Labs)
R
Ricardo Luna Gutierrez
Hewlett Packard Enterprise (Hewlett Packard Labs)
S
Soumyendu Sarkar
Hewlett Packard Enterprise (Hewlett Packard Labs)