Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation of existing speech deepfake detection methods in real-world scenarios, primarily caused by nonlinear distortions and time-frequency coupling introduced by acoustic front-end (AFE) processing. To mitigate this issue, the work presents the first systematic modeling of AFE’s impact on detection performance and proposes a time-frequency consistency learning framework. This framework employs an attention-driven soft alignment mechanism to alleviate temporal misalignment and incorporates a frequency-domain structural consistency constraint to counteract spectral distortions, thereby learning AFE-robust spoofing representations. Integrated with a simulated AFE pipeline encompassing echo cancellation, noise suppression, automatic gain control, and voice activity detection, the proposed method substantially reduces performance degradation and markedly enhances detection robustness under realistic conditions.
📝 Abstract
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.
Problem

Research questions and friction points this paper is trying to address.

speech deepfake detection
acoustic front-end
time-frequency distortion
robustness
real-world deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-Frequency Consistency Learning
Acoustic Front-End Robustness
Speech Deepfake Detection
Feature Invariance
Soft Temporal Alignment
🔎 Similar Papers
J
Jun Xue
School of Cyber Science and Engineering, Wuhan University
Z
Zhuolin Yi
School of Cyber Science and Engineering, Wuhan University
Y
Yanzhen Ren
School of Cyber Science and Engineering, Wuhan University
Y
Yihuan Huang
School of Cyber Science and Engineering, Wuhan University
J
Jiayu Xiong
Tongji university
Y
Yi Chai
School of Cyber Science and Engineering, Wuhan University
G
Guanxiang Feng
School of Cyber Science and Engineering, Wuhan University
J
Jiajun Liu
School of Cyber Science and Engineering, Wuhan University
Tong Zhang
Tong Zhang
Professor of GIS/Remote Sensing, Wuhan University
GeoAImachine learningtransport geography