π€ AI Summary
This study addresses the vulnerability of speech synthesis watermarking to model-driven reconstruction attacks, which undermines reliable source tracing. To this end, we propose Thrive, a framework for multi-bit generative watermarking in autoregressive text-to-speech (TTS) systems. Thrive introduces a novel adaptive multi-bit mechanism that eliminates reliance on a single carrier and supports both discrete token and continuous representation paradigms. Specifically, it synchronously injects intermediate representations via a Rise module and employs a Care module that integrates waveform and spectrogram experts for bit-level reliability selection. Experimental results demonstrate that Thrive maintains high synthesis fidelity while achieving an 87.6% watermark recovery accuracy under reconstruction attacks, thereby enabling identity tracing across tens of thousands of users.
π Abstract
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.