Score
Designs, builds, and evaluates algorithms and systems that detect manipulated or synthetically generated audiovisual content (images, video, and audio) by extracting and modeling forensic artifacts, temporal and spatial inconsistencies, and physiological or acoustic cues. Work includes constructing feature pipelines and training classifiers or anomaly detectors, curating and annotating evaluation datasets, measuring robustness to postprocessing and adversarial attacks, and producing forensic outputs for analysis or deployment.
Deepfakes—generated via diffusion models, GANs, and VAEs—increasingly produce high-fidelity multimodal synthetic content (e.g., face swapping, voice conversion, lip-sync), posing systemic threats to privacy, security, and democratic integrity. To address this, we present a systematic survey of state-of-the-art generation and detection techniques, introducing for the first time a unified theoretical framework that jointly models three dominant generative paradigms and their adversarial detection mechanisms, thereby exposing methodological bottlenecks in the “generation–detection” arms race. We further propose a generalizable cross-modal evaluation framework that rigorously benchmarks robustness, interpretability, and generalization capacity across approaches. Finally, we articulate a risk-tiered governance pathway grounded in technical feasibility and societal impact. Our work establishes foundational theory and methodology for designing robust, transferable multimodal deepfake detection systems.
The proliferation of generative AI has exacerbated the spread of deepfakes, while existing detection methods exhibit severe vulnerability to adversarial perturbations, limiting their practical utility against real-world threats. This work systematically surveys state-of-the-art deepfake detection under generative AI, focusing on two paradigms: fully synthetic content identification and spatiotemporal localization of authentic video manipulations. Centering on adversarial robustness as the primary evaluation criterion, we introduce an open-source, reproducible benchmark (GitHub) that integrates statistical anomaly analysis, hierarchical feature extraction, multimodal cue fusion—particularly visual artifacts and temporal inconsistencies—and advanced deep learning architectures. Empirical evaluation reveals that although current methods achieve high accuracy under benign conditions, they consistently degrade under minimal adversarial perturbations. To bridge the gap between algorithmic innovation and operational deployment, we propose design principles for robust, scalable, and multimodal detection frameworks resilient to adversarial interference.
The rapid advancement of generative AI—particularly GANs, diffusion models, and VAEs—has significantly intensified the risks and societal harms associated with synthetic imagery. While existing surveys predominantly focus on deepfake detection, they lack systematic coverage of multimodal digital forensics and emerging synthetic image identification techniques. To address this gap, we propose the first taxonomy of synthetic image detection methods explicitly designed for multimodal frameworks. Our survey comprehensively analyzes over 100 representative works published between 2019 and 2024, spanning key paradigms including frequency-domain analysis, texture anomaly modeling, neural artifact identification, cross-modal alignment, self-supervised pretraining, and large-model zero-shot discrimination. We consolidate more than ten mainstream public benchmarks into a structured knowledge graph, enabling rigorous algorithmic development, standardized benchmarking, and robustness evaluation. This work provides both theoretical foundations and practical guidance for advancing trustworthy multimodal forensic research.
AI-generated video detection methods suffer from poor generalization across diverse generative models. Method: This paper proposes a forensics-oriented frequency-domain enhancement approach that leverages wavelet decomposition to localize and replace critical frequency bands, thereby guiding the detector to focus on low-level, model-agnostic artifacts introduced by generators—rather than volatile high-level semantic inconsistencies. The method employs a lightweight classifier, single-source model training, and a multi-model generalization evaluation paradigm. Contribution/Results: Trained exclusively on videos synthesized by a single generator (e.g., SVD), the method achieves significantly higher cross-model detection accuracy than state-of-the-art methods on challenging benchmarks including NOVA and FLUX. This demonstrates that the learned features are highly robust and generalize effectively across unseen generative models, without requiring multi-source training data or architectural modifications.
The proliferation of AI-generated content (AIGC) has intensified risks including misinformation, copyright infringement, security vulnerabilities, and eroded public trust. To address these challenges, this paper proposes the first comprehensive multimodal AIGC detection framework for text, image, and audio modalities. It systematically integrates detection motivations, risk scenarios, and a taxonomy of technical approaches, prioritizing robustness, adaptability to evolving generative models, and human-AI collaborative verification. Key contributions include: (1) the first unified cross-modal detection methodology; (2) a dynamically adaptable detection paradigm capable of generalizing to novel generative models; and (3) formalization of a human-in-the-loop feedback loop as central to authenticity assurance. The framework synergizes observational analysis, linguistic statistical modeling, deep neural detectors, digital watermarking/fingerprinting, and ensemble learning. Empirical outcomes yield a practice-oriented guideline applicable across academic, journalistic, judicial, and industrial domains, while identifying critical challenges—adversarial perturbations, cross-domain generalization, and ethical boundaries—to inform policy formulation and tool development.
Existing video forensic methods typically target a single manipulation type (e.g., deepfakes or inpainting), rendering them inadequate for real-world scenarios where manipulation types are unknown and often co-occur. This paper introduces the first end-to-end, multi-purpose video forensic network capable of jointly detecting diverse manipulations—including deepfakes, inpainting, splicing, and editing—without prior knowledge of the manipulation type. Our method features a novel multi-scale hierarchical Transformer module that jointly models spatiotemporal anomalies and precisely localizes forged regions of arbitrary shape and size across scales. Additionally, it integrates multimodal forensic cues with multi-scale spatiotemporal features. Evaluated on a comprehensive multi-manipulation benchmark, our approach achieves state-of-the-art performance, while also matching or surpassing specialized detectors on single-type manipulation tasks—demonstrating significantly improved generalization and practical applicability.
To address the growing threat of video tampering in surveillance footage—which undermines its admissibility as judicial evidence—this paper presents a systematic survey of video forgery detection techniques tailored to security monitoring scenarios. We propose the first robustness evaluation framework specifically designed for real-world surveillance conditions, characterized by low resolution, high compression, and dynamic illumination variations. The framework integrates compression artifact analysis, temporal consistency verification, and hybrid feature extraction combining deep learning models (CNNs and LSTMs) with handcrafted features. For the first time, we conduct a comprehensive comparative analysis of three mainstream approaches—compression-based feature analysis, frame duplication detection, and machine learning–based methods—elucidating their respective applicability boundaries and performance limitations under practical surveillance constraints. Our empirical study identifies characteristic failure modes of existing detectors across typical surveillance conditions, thereby providing evidence-based guidance for forensic system design, algorithm optimization, and standardization efforts in digital video authentication.
This study addresses the challenge of detecting AI-manipulated visual evidence in judicial contexts and the lack of domain-specific benchmarks. To this end, it introduces the first courtroom-oriented benchmark and dataset for detecting AI-manipulated images. By integrating generative forgery techniques, image forensic analysis, and structured metadata annotation, this work establishes a novel evaluation framework encompassing multi-source evidence modalities, localized edits, and consumer-grade tool threat models. The authors release 1,505 samples, source code, and baseline models as open-source resources. Experimental evaluations reveal significant performance deficiencies of existing detectors in judicial evidence verification. Ultimately, this research provides critical support for advancing visual forgery detection within legal scenarios.
Existing deepfake detection methods rely heavily on artifacts left by generative models, leading to a significant drop in generalization when confronted with emerging generative architectures and interactive deception scenarios—such as video or voice impersonation—where the core threat lies in deceptive behavior rather than signal-level anomalies. This work breaks from conventional signal-centric paradigms by systematically integrating Speech Act Theory, Grice’s Cooperative Principle, and Cialdini’s Principles of Influence to construct a novel three-tiered analytical framework encompassing speech acts, dialogic interaction, and audience response. By introducing foundational social theories into media forensics, this framework not only exposes the “generalization illusion” inherent in current approaches but also establishes a new pathway for detecting deception in interactive deepfake contexts, while highlighting critical open challenges in the field.
This study addresses the scarcity of high-quality datasets for AI-generated video forgery detection and the difficulty of multimodal models in leveraging low-level cues for pixel-level localization. To this end, it constructs a unified forensic analysis framework encompassing forgery detection, artifact localization, and anomaly explanation. Methodologically, this work introduces ManiVid-38K, the first dataset featuring local manipulations with comprehensive annotations, and proposes a unified architecture based on multimodal large language models. Specifically, a forensic evidence router is designed to share low-level features, while a prompt distillation module injects spatial priors. Experimental results demonstrate that the proposed method maintains detection accuracy comparable to specialized classifiers while significantly enhancing performance in artifact localization and anomaly explanation.
This work addresses the limitations of existing audiovisual AIGC detection methods, which rely on the assumption of audio-visual consistency and often degrade in generalizable scenarios. To overcome this, we propose DAV-Det, the first modality-disentangled detection framework tailored for general-purpose AIGC content. Departing from conventional feature-level fusion, DAV-Det adopts decision-level fusion to independently model forgery traces in audio and video modalities. The visual branch leverages a three-level granularity representation—encompassing global, patch, and clip-level features—while the audio branch employs a gated dual-branch architecture operating in both time and frequency domains to capture anomalies. Evaluated on the IJCAI-ECAI 2026 DDL 2.0 Workshop challenge for general AIGC audiovisual detection, our method achieves state-of-the-art performance with a score of 0.8460, demonstrating superior robustness and generalization.
研究提出TRACE框架,通过预训练视频模型的速度响应提取表示并建模帧间一致性,有效检测AI生成的视频,超越现有方法。