🤖 AI Summary
This study addresses the limited generalization capability of existing motion-based AI-generated video detection methods, which rely heavily on inter-frame motion biases present in datasets rather than genuine forgery artifacts. The authors systematically evaluate four state-of-the-art motion-based detectors and, for the first time, explicitly demonstrate that these models exploit motion shortcuts through biased training and testing data. Through data rebalancing, spatial augmentations, and cross-dataset evaluations, they reveal the fragility of such approaches. Experiments show that when motion bias is eliminated in a newly curated dataset, all motion-based detectors degrade to random-chance performance, whereas frequency-domain detectors maintain consistently high accuracy, underscoring the superior robustness of frequency-based methods.
📝 Abstract
The visual quality of AI-generated videos has improved drastically in recent years, making it increasingly difficult for humans to distinguish between real and synthetic media. In this work, we evaluate the robustness and applicability of four state-of-the-art motion-based AI-generated video detectors. We identify significant preprocessing and sampling biases in these methods and demonstrate that they account for a substantial portion of their reported performance. Furthermore, we find that these detectors are highly sensitive to motion patterns specific to their evaluation datasets, where AI-generated videos generally exhibit less inter-frame movement than real videos. We show that for all detectors, performance collapses to near-random levels when evaluated on a dataset that does not contain this motion bias. Additionally, through dataset rebalancing and the application of simple spatial augmentations, we observe severe performance degradation across all evaluated models. In contrast, we find that an existing frequency-based detector maintains strong performance across all evaluated datasets, suggesting that frequency-based approaches may offer a more generalizable path forward for AI-generated video detection. We hope that our work raises awareness towards these vulnerabilities and encourages the development of more representative, unbiased datasets and more robust evaluation protocols.