🤖 AI Summary
This study addresses the vulnerability of audio deepfake detectors to adversarial perturbations and the inherent difficulty existing purification methods face in balancing noise removal with the preservation of forgery cues. To overcome these limitations, this work proposes the DGAP framework, which introduces a novel detection-feedback-driven adaptive purification mechanism. Specifically, the framework leverages detection score shifts as a reference-free metric to dynamically modulate purification intensity via diffusion models, enabling the discrimination between benign and adversarial samples without requiring detector retraining. The proposed approach achieves state-of-the-art defensive performance across diverse attack scenarios while ensuring lossless processing for benign inputs and demonstrating robust resilience against adaptive attacks.
📝 Abstract
Audio deepfake detectors remain vulnerable to adversarial perturbations that suppress the acoustic cues used for detection, allowing manipulated utterances to evade the detector. Although existing defenses can improve robustness, they require retraining the detector or introduce additional distortion. Diffusion-based purification instead leaves the pretrained detector unchanged, but existing methods use the same purification strength for all inputs, creating a trade-off between removing adversarial perturbations and preserving the subtle spoofing cues needed for detection. In this paper, we propose Detection-Guided Adaptive Purification (DGAP), a diffusion-based defense that adjusts purification strength per input. Building on the observation that a light purification perturbs the detector score of an adversarial input far more than that of a benign one, the framework uses the resulting score shift as a reference-free indicator of adversarial manipulation. Inputs with small shifts are passed unchanged, whereas flagged inputs undergo stronger purification before final detection. We evaluate the framework against three adversarial attack settings across three deepfake detectors, and compare it with nine existing defenses. Our results show that DGAP achieves the best defense performance across all detectors while leaving benign inputs nearly unaffected, and remains effective under the defense-aware adaptive attack.