ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of high-quality datasets for AI-generated video forgery detection and the difficulty of multimodal models in leveraging low-level cues for pixel-level localization. To this end, it constructs a unified forensic analysis framework encompassing forgery detection, artifact localization, and anomaly explanation. Methodologically, this work introduces ManiVid-38K, the first dataset featuring local manipulations with comprehensive annotations, and proposes a unified architecture based on multimodal large language models. Specifically, a forensic evidence router is designed to share low-level features, while a prompt distillation module injects spatial priors. Experimental results demonstrate that the proposed method maintains detection accuracy comparable to specialized classifiers while significantly enhancing performance in artifact localization and anomaly explanation.
📝 Abstract
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% J&F) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).
Problem

Research questions and friction points this paper is trying to address.

video forensics
manipulated video
forgery detection
artifact grounding
anomaly explanation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Forensics
Multimodal Large Language Models
Artifact Grounding
Explainable AI
Manipulated Video Detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.