MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited generalizability of existing AI-generated video detection methods, which often overlook textual semantics and fine-grained temporal patterns. To this end, we propose MTOR, a novel framework that uniquely integrates textual semantic complementarity with multi-level temporal stability for detection. Specifically, MTOR fuses visual-textual multimodal semantics and models Temporal Over-Regularity (TOR) across coarse, medium, and fine granularities to precisely identify AI-generated videos. Extensive experiments demonstrate that MTOR outperforms 16 baseline methods across five benchmarks and exhibits strong robustness against 12 real-world perturbations, achieving state-of-the-art performance.
πŸ“ Abstract
The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.
Problem

Research questions and friction points this paper is trying to address.

AI-generated video detection
generalizability
multimodal semantics
temporal regularity
synthetic video
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Semantics
Temporal Over-Regularity
AI-Generated Video Detection
Generalizability
Fine-grained Visual Representations
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
H
Hang Wang
Xi’an Jiaotong University, Xi’an, China, and The Hong Kong Polytechnic University, Hong Kong, China
Chao Shen
Chao Shen
Chair Professor, Xi'an Jiaotong University
AI SecuritySoftware SecurityControl System
L
Lei Zhang
The Hong Kong Polytechnic University, Hong Kong, China
Zhi-Qi Cheng
Zhi-Qi Cheng
Assistant Professor @ UW | Graduate Faculty | Ex-CMU, Google, Microsoft | Intel & IBM PhD Fellowship
multimedia processingmultimedia understandingmultimodal foundation model