🤖 AI Summary
This study addresses the vulnerability and insufficient robustness of conventional post-hoc speech watermarking by proposing a TTS-native watermarking framework. The method pioneers the utilization of the complete latent space of neural audio codecs for multi-layer information embedding, completing the marking process prior to waveform decoding. Furthermore, the pre-trained model remains frozen while only the watermarking modules are jointly optimized. Experimental results demonstrate that the proposed framework maintains high detection rates under various digital signal processing and re-synthesis attacks, achieving full-pipeline traceability. Meanwhile, the naturalness and quality of the synthesized speech remain comparable to the original level.
📝 Abstract
Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.