🤖 AI Summary
This study addresses the challenge of deploying high-parameter self-supervised learning models for audio deepfake detection on resource-constrained devices by proposing a task-aware joint pruning and distillation framework. The method introduces the first forgery-discriminative joint optimization mechanism, overcoming the performance bottlenecks inherent in conventional content-centric compression approaches. Efficient model compression is achieved through the synergy of cross-domain knowledge distillation and motion-guided structured pruning. Experimental results demonstrate that the compressed model contains only 31.9M parameters with a 6.3-fold reduction in computational cost, while the average performance across multiple datasets degrades by merely 1.3%. By preserving strong generalization capabilities, the proposed approach exhibits significant potential for edge deployment.
📝 Abstract
Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parameters) prevent deployment on resource-constrained devices. Existing compression methods, designed mainly for content-centric tasks, struggle to maintain competitive performance when directly adapted to deepfake detection. We propose a Task-Aware Joint Pruning and Distillation framework that combines cross-domain knowledge distillation with movement-guided structured pruning to transfer forgery-discriminative knowledge and preserve critical structures under aggressive compression. Our framework reduces the model to 31.9M parameters with 6.3$\times$ FLOPs reduction, with an average performance drop of only 1.30\% across multiple datasets compared to the uncompressed baseline, demonstrating strong potential for on-device deployment.