🤖 AI Summary
This study addresses the insufficient robustness of existing neural audio watermarking against speech enhancement (SE) attacks by proposing a black-box watermark removal framework. This framework cascades Gaussian noise with SE models, revealing that generative SE severely disrupts watermarks through the reconstruction of harmonic regions. By integrating both discriminative and generative SE techniques, we conduct adversarial evaluations on six mainstream watermarking methods, including AudioSeal. Experimental results demonstrate that the proposed attack significantly outperforms existing resynthesis approaches, confirming that SE poses a critical threat to current watermark security. Consequently, this work calls for establishing an SE-aware robustness evaluation paradigm to better safeguard neural audio watermarking systems.
📝 Abstract
Neural audio watermarks are increasingly deployed in commercial speech generation systems to make AI-generated speech traceable, yet their robustness has been studied mainly under conventional signal distortions. Since a watermark can be regarded as imperceptible noise added to the speech signal, a natural question is whether speech enhancement (SE), as a denoising model, can remove it. In this paper, we cascade Gaussian noise with SE models as a black-box watermark removal attack, covering both discriminative and generative SE paradigms, against six neural watermarks: AudioSeal, WavMark, SilentCipher, Timbre, Perth, and AlignMark. Experimental results show that the proposed attack significantly outperforms existing neural re-synthesis methods in watermark removal. In particular, we find that generative SE, which reconstructs the harmonic regions of speech while denoising, is highly destructive to watermarks. These findings show that SE poses a serious threat to current audio watermarking methods, and we call for SE-aware robustness evaluation in watermark design.