GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过GenTraceBench评估了音频深度伪造在模型适应前后的指纹一致性问题,采用多种训练和测试方法发现不同适应策略对指纹的影响。
📝 Abstract
Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain valid after adaptation? We present GenTraceBench, a controlled benchmark spanning five TTS architectures, 16 pre-/post-training variants, and 49,728 utterances generated with fixed texts and speaker prompts. Under a train-on-foundation, test-on-adapted protocol, we evaluate binary detection, closed-set attribution, and open-set verification. DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones. Repeated training runs confirm the largest W2V-BERT attribution drop, while a data-mixture control with comparable speech quality shows that composition change need not cause drift. In W2V-BERT verification, multi-shot enrollment reduces EER for the SFT condition from 44.4% to 11.0%, whereas the SingNet-only condition remains at or above 45% EER.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake forensics
TTS systems
fingerprint preservation
adaptation
supervised fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

GenTraceBench
TTS architectures
fingerprint preservation
supervised fine-tuning (SFT)
error rate reduction
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25