🤖 AI Summary
This study addresses the limitations of existing video safety detection methods, which often overlook multi-label characteristics and lack controllable trade-offs between precision and recall. To tackle these challenges, this work proposes ATPO, a reinforcement learning framework that integrates vision-language models with an adaptive Tversky reward mechanism. By dynamically adjusting penalties for false positives and false negatives, the proposed approach enables controllable operating points, effectively resolving the complexities of multi-label video safety detection. Experimental results demonstrate that this method substantially improves the Jaccard index from 40.66 to 75.44 on the SafeWatch-Bench benchmark, validating its superior performance and practical utility in complex scenarios.
📝 Abstract
The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .