🤖 AI Summary
This study addresses the generalization bottlenecks in audio deepfake detection caused by domain shift and gradient conflicts inherent in meta-learning. To overcome these challenges, we propose Domain Gradient Surgery (DGS). The core innovation lies in introducing an asymmetric gradient projection strategy that eliminates conflicts between meta-training and meta-testing phases, alongside a layer-wise dynamic intervention mechanism (LW-DGS) designed to achieve conflict-free optimization trajectories. Experimental results demonstrate that the proposed method reduces the average relative Equal Error Rate (EER) by 5.29% and 4.04%, respectively, across standard benchmarks. These improvements significantly enhance detection robustness in cross-domain scenarios, establishing DGS as an effective solution for mitigating gradient interference during meta-optimization in audio spoofing detection tasks.
📝 Abstract
Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.