Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction
This study addresses the vulnerability of speech synthesis watermarking to model-driven reconstruction attacks, which undermines reliable source tracing. To this end, we propose Thrive, a framework for multi-bit generative watermarking in autoregressive text-to-speech (TTS) systems. Thrive introduces a novel adaptive multi-bit mechanism that eliminates reliance on a single carrier and supports both discrete token and continuous representation paradigms. Specifically, it synchronously injects intermediate representations via a Rise module and employs a Care module that integrates waveform and spectrogram experts for bit-level reliability selection. Experimental results demonstrate that Thrive maintains high synthesis fidelity while achieving an 87.6% watermark recovery accuracy under reconstruction attacks, thereby enabling identity tracing across tens of thousands of users.