Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of gating networks to unseen attacks in audio deepfake detection, caused by their neglect of acoustic temporal cues. To overcome this, we propose a domain-adaptive dual-gated Mixture-of-Experts (MoE) framework. Methodologically, the approach fuses raw waveforms with self-supervised learning (SSL) features and employs Sinc-layer filters to extract multi-granularity signals. Furthermore, it introduces an innovative dual-gating mechanism incorporating domain prototypes to enable intelligent expert routing based on implicit forgery patterns. Experimental results demonstrate that the proposed model significantly outperforms baselines on challenging datasets, achieving a 40.8% relative reduction in equal error rate. These findings highlight its superior cross-domain generalization capability and architectural efficiency.
📝 Abstract
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.
Problem

Research questions and friction points this paper is trying to address.

Speech Deepfake Detection
Domain Adaptation
Generalization
Mixture of Experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Experts
Dual-Gating Mechanism
Domain Adaptation
Speech Deepfake Detection
Self-Supervised Learning
🔎 Similar Papers
S
Siqing Qin
Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University
Z
Zhe Li
Speech, Language, and Cognition Laboratory, The University of Hong Kong
Kong Aik Lee
Kong Aik Lee
The Hong Kong Polytechnic University, Hong Kong
Speaker and Spoken Language RecognitionSpeech ProcessingDigital Signal ProcessingSubband
M
Man-Wai Mak
Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University