GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization capability of deepfake speech detection and the frequent neglect of fine-grained artifacts by proposing the GLAD detector. Methodologically, it presents a pioneering empirical analysis revealing inherent deficiencies in self-supervised learning (SSL) models, which motivates the design of a hierarchical global-local backbone network to fuse multi-granularity features. An adaptive gating mechanism is further introduced to dynamically optimize inter-layer weights. Additionally, a composite data augmentation strategy termed SaniBoost is proposed to effectively mitigate domain shift and environmental noise interference. Experimental results demonstrate that GLAD significantly outperforms existing state-of-the-art methods, exhibiting superior robustness in cross-domain, unseen scenarios. The source code will be made publicly available upon publication.
📝 Abstract
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.
Problem

Research questions and friction points this paper is trying to address.

Speech Deepfake Detection
Self-Supervised Learning
Domain Generalization
Fine-grained Artifacts
Out-of-Distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Deepfake Detection
Global-Local Adaptive Detector
Hierarchical Adaptive Gating
Self-Supervised Learning
Data Augmentation
🔎 Similar Papers