Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video backbone networks struggle to detect temporal artifacts—such as inter-frame inconsistencies in AI-generated videos—due to excessive aggregation of spatiotemporal information by global readout layers. This work is the first to reveal how this mechanism suppresses detection performance and introduces V-PVP, a lightweight, plug-and-play readout module. By replacing conventional global aggregation with a dual-stream gated architecture built upon patch velocity fields, V-PVP preserves local temporal dynamics and cross-patch relationships with negligible parameter overhead, effectively reactivating the temporal awareness of frozen backbones. The method supports both end-to-end fine-tuning and linear probing, achieving a 95.28 AUC on AIGVDBench; notably, merely swapping the readout layer enables frozen video backbones to substantially outperform image-pretrained models.
📝 Abstract
AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In practice, however, video backbones with standard global readouts often fail to outperform strong image pretrained probes on AIGV benchmarks. We attribute this gap to excessive spatiotemporal aggregation in the readout. Video pretrained backbones tend to compress each frame into a single global descriptor. This compression suppresses local patch level temporal dynamics and discards inter patch relations, which are precisely the cues that AIGV detection most reliably depends on. Based on this, we propose Velocity Gated Patch Velocity Profiling (V-PVP), a lightweight readout that replaces only the aggregation layer with two parallel streams over the patch velocity field, adding only about $0.5$M trainable parameters. V-PVP serves as a general plug-and-play module that consistently improves performance across diverse video backbones under both end-to-end fine-tuning and linear probing settings. Our method reaches \textbf{95.28} AUC on AIGVDBench while keeping the backbone fully frozen. The results show that simply replacing the aggregation layer reactivates the temporal potential of frozen video backbones, restoring their advantage on AIGV detection. Code is available at https://anonymous.4open.science/r/PVP-81B3/.
Problem

Research questions and friction points this paper is trying to address.

AI-generated video detection
temporal artifacts
video backbones
spatiotemporal aggregation
inter-frame inconsistencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

video backbone
temporal artifacts
patch velocity profiling
readout module
AI-generated video detection
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
M
Manni Cui
Huazhong University of Science and Technology
Z
Ziheng Qin
Institute of Automation, Chinese Academy of Sciences
Z
ZiAn Wang
Jilin University
Ruiqi Liu
Ruiqi Liu
Texas Tech University
nonparametric methodsmachine learningeconometrics
D
Dianyuan Zou
Huazhong University of Science and Technology
J
Jianglan Wei
Huazhong University of Science and Technology
H
Han Zhou
Huazhong University of Science and Technology
Y
Yu Liu
Huazhong University of Science and Technology
J
Jingrui Xu
Huazhong University of Science and Technology
W
Wenhao Wang
Vast Intelligence Lab
Z
Zhenyu Zhang
Huazhong University of Science and Technology