SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing audio watermarking methods predominantly embed information at the signal level, rendering them vulnerable to severe distortions and inadequate for ensuring traceability and verifiability of synthetic speech. This work proposes the first training-free, semantic-level voice watermarking framework, which encodes watermark information into the semantic content conveyed by the generated speech and detects it via automatic speech recognition followed by natural language processing on the transcribed text. By elevating watermark embedding from the signal domain to the semantic domain, the method substantially enhances robustness against diverse editing attacks, including signal processing and compression. Evaluated on the AudioMarkBench benchmark, the proposed approach achieves detection rates that surpass the current state-of-the-art by 4.6% and 16.9% under signal processing and compression attacks, respectively, at a 1% false positive rate.
📝 Abstract
As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow. Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1\%, SSTMark improves the average detection rate by 4.6\% and 16.9\% on signal-processing edits and compression edits, respectively.
Problem

Research questions and friction points this paper is trying to address.

speech watermarking
synthetic speech
robustness
semantic-level
attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic-level watermarking
training-free
speech watermarking
text watermarking
robustness