IVT-Guard: All-in-One Reasoning Model for AI-Generated Content Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing AIGC detectors, including modality singularity, insufficient reasoning capabilities, and the scarcity of multimodal data, by constructing the IVT-Set dataset and proposing the IVT-Guard framework. Methodologically, it introduces a pioneering three-stage training paradigm that combines artifact-aware injection with an evidence-verdict consistency strategy to resolve optimization conflicts between reasoning and detection. By integrating multi-granularity chain-of-thought, supervised fine-tuning, and group relative policy optimization, the framework achieves unified and interpretable detection across images, videos, and text. Experimental results demonstrate that the proposed approach attains state-of-the-art performance in in-domain, out-of-domain, and cross-dataset scenarios while providing faithful and reliable reasoning processes.
📝 Abstract
The rapid proliferation of highly realistic AI-Generated Content (AIGC) necessitates robust and interpretable detection mechanisms. However, existing detectors are predominantly confined to single modalities and provide binary outputs without reasoning. While Multimodal Large Language Models (MLLMs) present a promising solution, their development is constrained by the scarcity of multimodal reasoning data and the reasoning-detection optimization dilemma, where explicit reasoning supervision can compromise detection accuracy. To this end, we introduce IVT-Set, a comprehensive dataset comprising over 152K diverse image, video, and text samples equipped with multi-granularity Chain-of-Thought (CoT) reasoning trajectories. Based on it, we propose IVT-Guard, a pioneering framework for unified and interpretable AIGC detection across image, video, and text modalities. Furthermore, to overcome the aforementioned optimization dilemma, we design a novel three-stage training paradigm: Artifact-Aware Pre-training, Artifact-to-Evidence Supervised Fine-Tuning via artifact-aware injection, and Evidence-Verdict Consistency Group Relative Policy Optimization. Extensive experiments demonstrate that IVT-Guard achieves state-of-the-art detection performance across in-domain, out-of-domain, and cross-dataset settings while delivering faithful reasoning. Code and data will be released.
Problem

Research questions and friction points this paper is trying to address.

AI-Generated Content Detection
Multimodal Reasoning
Interpretability
Optimization Dilemma
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-Generated Content Detection
Multimodal Reasoning
Chain-of-Thought
Group Relative Policy Optimization
Three-stage Training Paradigm
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
Hongwei Niu
Hongwei Niu
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Xiamen 361005, P.R. China
Yunpeng Luo
Yunpeng Luo
Bytedance Inc.
H
Hanjun Li
Tencent YouTu Lab
Z
Ziyin Zhou
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Xiamen 361005, P.R. China
Jianghang Lin
Jianghang Lin
Xiamen University
Multimodal Large Language ModelVision-Language ModelSemi/Weakly-Supervised Learning
Ke Yan
Ke Yan
Tencent
Multimodal LLMComputer VisionImage/Video Understanding
S
Shouhong Ding
Tencent YouTu Lab
Shengchuan Zhang
Shengchuan Zhang
Xiamen University
computer visionmachine learning
L
Liujuan Cao
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Xiamen 361005, P.R. China