๐ค AI Summary
This study addresses the limitations of existing DeepFake datasets, including linguistic homogeneity, the absence of audio modalities, and insufficient consent compliance, by constructing the first informed-consent multilingual audio-visual manipulation benchmark. Comprising 399,000 clips across five languages, this benchmark integrates eleven video forgery methods and four voice cloning engines through a modular pipeline. Our investigation reveals that detection difficulty is highly dependent on generation pairing strategies, and demonstrates that retaining authentic audio leads to a significant degradation in detection performance. Furthermore, we identify a pronounced misalignment between human perceptual judgment and machine-based detection capabilities. This work provides a comprehensive, ethically grounded resource for advancing multimodal deepfake detection research.
๐ Abstract
Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.