CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing deepfake datasets, which predominantly feature benign content generated by outdated models and thus fail to adequately evaluate highly realistic forgery threats in surveillance scenarios. To this end, we construct CCDF, a benchmark dataset comprising 1,840 surveillance videos synthesized using state-of-the-art commercial generative systems such as Sora 2 and VEO 3.1. Through metadata standardization and multi-stage data cleaning, potential detection shortcuts are systematically eliminated. Experimental evaluations of ten state-of-the-art detectors demonstrate significant performance degradation on the proposed CCDF benchmark. These findings reveal critical shortcomings in current evaluation protocols and establish a more realistic assessment platform for detecting high-fidelity deepfakes in surveillance contexts.
📝 Abstract
Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF's generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.
Problem

Research questions and friction points this paper is trying to address.

Deepfake Detection
Surveillance Footage
Benchmark Dataset
Generative AI
Video Forensics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deepfake Detection
Benchmark Dataset
Surveillance Footage
Generative AI
Video Forensics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.