What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of cross-domain generalization in deepfake speech detection caused by neural codec-based speech synthesis, identifying an urgent need to determine which acoustic representations remain discriminative under shifting generative mechanisms. By comparing various acoustic representations, we find that unvoiced residual statistics exhibit exceptional robustness against unseen codecs. Leveraging their complementarity with XLS-R features, we propose MN-P, a dual-view detector that fuses hierarchical XLS-R representations with pooled residual statistics via adaptive gating and employs a low-capacity linear classifier. The proposed approach achieves a 54.2% relative reduction in overall equal error rate (EER) and a 60.9% reduction on unseen codecs, significantly outperforming retrained state-of-the-art systems.
📝 Abstract
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.
Problem

Research questions and friction points this paper is trying to address.

speech deepfake detection
neural codec
generalization
acoustic representations
codec shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Deepfake Detection
Neural Codec
No-Vocals Residuals
Dual-View Detector
Adaptive Gating
🔎 Similar Papers
No similar papers found.
J
Jiajun Xu
Department of Engineering Physics, Tsinghua University, Beijing, China
Menglu Li
Menglu Li
Toronto Metropolitan University
Audio ProcessingDeep Learning
X
Xiao-Ping Zhang
Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China