Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of speech synthesis watermarking to model-driven reconstruction attacks, which undermines reliable source tracing. To this end, we propose Thrive, a framework for multi-bit generative watermarking in autoregressive text-to-speech (TTS) systems. Thrive introduces a novel adaptive multi-bit mechanism that eliminates reliance on a single carrier and supports both discrete token and continuous representation paradigms. Specifically, it synchronously injects intermediate representations via a Rise module and employs a Care module that integrates waveform and spectrogram experts for bit-level reliability selection. Experimental results demonstrate that Thrive maintains high synthesis fidelity while achieving an 87.6% watermark recovery accuracy under reconstruction attacks, thereby enabling identity tracing across tens of thousands of users.
πŸ“ Abstract
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.
Problem

Research questions and friction points this paper is trying to address.

speech watermarking
text-to-speech
reconstruction attacks
source attribution
generative watermarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Watermarking
Speech Synthesis
Autoregressive TTS
Reconstruction Attacks
Source Attribution
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Weizhi Liu
Weizhi Liu
εŽδΈœεΈˆθŒƒε€§ε­¦
AIGC securityGenerative watermarking
Y
Yue Li
Huaqiao University, Xiamen, China
H
Hui Tian
Huaqiao University, Xiamen, China
Z
Zhaoxia Yin
East China Normal University, Shanghai, China