🤖 AI Summary
This study addresses the limitations of existing audio watermarking attacks, which typically rely on model feedback, incur high costs, and degrade audio quality. We propose a query-free black-box watermark removal method that, for the first time, exploits subtle non-speech artifacts introduced during embedding as attack cues. By learning diverse artifacts to extract complementary feature patterns and incorporating an adaptive scaling mechanism, our approach amplifies and suppresses watermark signals under perceptual quality constraints, overcoming the limitations of conventional generative reconstruction. Experiments demonstrate that the proposed method achieves average attack success rates of 0.92 and 0.96 across two datasets and four mainstream watermarking schemes, while significantly outperforming existing baselines in preserving audio quality.
📝 Abstract
Audio watermarking protects digital speech by embedding imperceptible signals for ownership verification and misuse tracing. However, the security of learning-based watermarking remains insufficiently understood under realistic adversarial removal, where attackers cannot access or query the watermark encoder, decoder, or detector. Existing attacks either rely on model feedback, require clean-watermarked pairs, or reconstruct the waveform with generative models, often leading to high query costs, limited generalization, or degraded perceptual quality. In this paper, we propose DeMark, a query-free black-box attack for quality-preserving audio watermark removal. Our key insight is that watermark embedding, while perceptually hidden, can introduce subtle non-speech artifacts in the time-frequency domain that are not fully aligned with natural speech. DeMark removes watermarks by suppressing these artifacts through two stages: Diverse Artifact Learning, which extracts complementary non-stationary and stationary artifact patterns, and Adaptive Artifact Scaling, which adaptively combines and amplifies them under quality-preserving constraints. Across two speech datasets and four state-of-the-art watermarking methods, DeMark achieves average attack success rates of 0.92 and 0.96 while consistently preserving higher perceptual quality than existing adaptive attacks. These results reveal a practical vulnerability of current audio watermarking systems and call for more robust watermark designs against query-free adversarial removal.