Tracing and Relearning Detection Evidence in Text-to-Speech Systems

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the originating stages of deepfake detection evidence within text-to-speech (TTS) pipelines. Leveraging the F5-TTS and BigVGAN architectures, controlled resynthesis experiments identify the acoustic model as the primary source of such evidence. The research reveals that adversarial fine-tuning can substantially degrade the performance of fixed detectors, whereas detector adaptation strategies effectively restore generalization capabilities. The principal contributions lie in elucidating how acoustic model updates interfere with detection evidence and in validating the robustness of adaptive training against unseen generated outputs. Experimental results demonstrate that the proposed approach reduces the equal error rate (EER) on LibriSpeech from 19.42% to 7.46%, significantly enhancing the detection of previously unseen F5-TTS variants.
📝 Abstract
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake detection
text-to-speech
detection evidence tracing
acoustic model
detector adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Detection Evidence Tracing
Adversarial Fine-tuning
Detector Adaptation
Text-to-Speech Deepfake
Acoustic Model
🔎 Similar Papers
No similar papers found.