Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake Videos

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of motion-aware deepfake video detection and attribution by constructing MAD, the first dedicated benchmark for this task, and proposing the MoDA defense framework. By integrating spatial semantic and frequency-domain features through cross-domain alignment and multi-scale aggregation mechanisms, MoDA precisely captures high-frequency artifacts at motion boundaries and spectral fingerprints to enable effective identification and source tracing. Experimental results demonstrate that the proposed method achieves an in-distribution detection accuracy of 94.8%, improves cross-dataset generalization performance by 10%–25%, and attains an attribution accuracy of 91.5%, exhibiting superior zero-shot transfer capabilities.
📝 Abstract
Pose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers. To better understand this emerging threat, we construct the first MAD-specific benchmark and measurement framework, containing over 1.5 million frames that mix 1,363 real and 30,122 synthetic videos from six controllable generators, with realistic perturbations and open-world evaluation splits. Then, we dissect MAD and discover that, despite their global coherence, these videos betray faint yet reliable cues: because the model relies on limited input frames for motion synthesis, it must predict and simulate coherent movement at motion boundaries, thereby producing high-frequency artifacts along with model-specific spectral fingerprints. Based on the observations obtained from analysis on dataset, we propose MoDA, the first defense framework tailored to detect and attribute MAD videos. MoDA couples spatial semantics with steganalysis-rich frequency features via cross-domain alignment and multi-scale aggregation, achieving 94.8% in-distribution and 89.1% cross-dataset detection accuracy gains of 10% to 25% over prior work and 91.5% model attribution accuracy. MoDA achieves 81.94% accuracy on 200 clips produced by two unseen commercial MAD platforms, indicating promising zero-shot transfer, and 78.13% detection accuracy on 1,200 unseen MAD video clips (55k frames in total) collected from the open Internet. Under white-box, gray-box, and black-box adaptive attacks, MoDA maintains relatively stable detection and attribution performance while the accuracies of the baselines drop rapidly.
Problem

Research questions and friction points this paper is trying to address.

Motion-Aware Deepfake
Deepfake Detection
Source Attribution
Pose-guided Diffusion Models
Video Forensics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion-Aware Deepfake
Cross-domain Alignment
Frequency Features
Video Detection and Attribution
Robustness
🔎 Similar Papers
No similar papers found.