BabelFake: A Multilingual Audio-Visual DeepFake Benchmark

๐Ÿ“… 2026-10-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitations of existing DeepFake datasets, including linguistic homogeneity, the absence of audio modalities, and insufficient consent compliance, by constructing the first informed-consent multilingual audio-visual manipulation benchmark. Comprising 399,000 clips across five languages, this benchmark integrates eleven video forgery methods and four voice cloning engines through a modular pipeline. Our investigation reveals that detection difficulty is highly dependent on generation pairing strategies, and demonstrates that retaining authentic audio leads to a significant degradation in detection performance. Furthermore, we identify a pronounced misalignment between human perceptual judgment and machine-based detection capabilities. This work provides a comprehensive, ethically grounded resource for advancing multimodal deepfake detection research.
๐Ÿ“ Abstract
Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.
Problem

Research questions and friction points this paper is trying to address.

Audio-Visual DeepFake
Multilingual Benchmark
DeepFake Detection
Voice Cloning
Face Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Benchmark
Audio-Visual DeepFake
Modular Generation Pipeline
Voice Cloning
Cross-language Evaluation
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
C
Carlotta Segna
TU Darmstadt & Hessian.AI, Germany
J
Joel Tschesche
TU Darmstadt & Hessian.AI, Germany
Anna Rohrbach
Anna Rohrbach
Professor, TU Darmstadt, Germany
Vision and LanguageArtificial IntelligenceMultimodal Grounded Learning