🤖 AI Summary
The proliferation of generative AI has exacerbated the spread of deepfakes, while existing detection methods exhibit severe vulnerability to adversarial perturbations, limiting their practical utility against real-world threats. This work systematically surveys state-of-the-art deepfake detection under generative AI, focusing on two paradigms: fully synthetic content identification and spatiotemporal localization of authentic video manipulations. Centering on adversarial robustness as the primary evaluation criterion, we introduce an open-source, reproducible benchmark (GitHub) that integrates statistical anomaly analysis, hierarchical feature extraction, multimodal cue fusion—particularly visual artifacts and temporal inconsistencies—and advanced deep learning architectures. Empirical evaluation reveals that although current methods achieve high accuracy under benign conditions, they consistently degrade under minimal adversarial perturbations. To bridge the gap between algorithmic innovation and operational deployment, we propose design principles for robust, scalable, and multimodal detection frameworks resilient to adversarial interference.
📝 Abstract
The rapid advancement of Generative Artificial Intelligence has fueled deepfake proliferation-synthetic media encompassing fully generated content and subtly edited authentic material-posing challenges to digital security, misinformation mitigation, and identity preservation. This systematic review evaluates state-of-the-art deepfake detection methodologies, emphasizing reproducible implementations for transparency and validation. We delineate two core paradigms: (1) detection of fully synthetic media leveraging statistical anomalies and hierarchical feature extraction, and (2) localization of manipulated regions within authentic content employing multi-modal cues such as visual artifacts and temporal inconsistencies. These approaches, spanning uni-modal and multi-modal frameworks, demonstrate notable precision and adaptability in controlled settings, effectively identifying manipulations through advanced learning techniques and cross-modal fusion. However, comprehensive assessment reveals insufficient evaluation of adversarial robustness across both paradigms. Current methods exhibit vulnerability to adversarial perturbations-subtle alterations designed to evade detection-undermining reliability in real-world adversarial contexts. This gap highlights critical disconnect between methodological development and evolving threat landscapes. To address this, we contribute a curated GitHub repository aggregating open-source implementations, enabling replication and testing. Our findings emphasize urgent need for future work prioritizing adversarial resilience, advocating scalable, modality-agnostic architectures capable of withstanding sophisticated manipulations. This review synthesizes strengths and shortcomings of contemporary deepfake detection while charting paths toward robust trustworthy systems.