🤖 AI Summary
This study addresses the challenge of cross-domain generalization in deepfake speech detection caused by neural codec-based speech synthesis, identifying an urgent need to determine which acoustic representations remain discriminative under shifting generative mechanisms. By comparing various acoustic representations, we find that unvoiced residual statistics exhibit exceptional robustness against unseen codecs. Leveraging their complementarity with XLS-R features, we propose MN-P, a dual-view detector that fuses hierarchical XLS-R representations with pooled residual statistics via adaptive gating and employs a low-capacity linear classifier. The proposed approach achieves a 54.2% relative reduction in overall equal error rate (EER) and a 60.9% reduction on unseen codecs, significantly outperforming retrained state-of-the-art systems.
📝 Abstract
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.