🤖 AI Summary
This study addresses the inherent trade-off between security and embedding rate in multi-bit text watermarking by proposing an f-divergence-constrained framework. The proposed approach unifies general f-divergence-based security definitions, encompassing both total variation and Kullback-Leibler divergences, and optimizes token distributions through a tailored coding scheme to provide rigorous security guarantees for each key. Experimental results demonstrate that the method significantly reduces watermark detectability while maintaining competitive message recovery performance and text generation quality. Consequently, this work achieves efficient information embedding and secure decoding with minimal perceptibility, offering a principled solution for balancing robustness and stealth in multi-bit text watermarking systems.
📝 Abstract
We introduce a framework for multi-bit text watermarking with security defined directly through $f$-divergence from the base language model distribution. Unlike prior approaches that focus on average-key distortion-freeness or a particular statistical distance, our formulation supports general $f$-divergences, including total variation and KL divergence, and enforces the guarantee for each realized key and embedded message. We develop a coding-based watermarking scheme that optimally biases next-token distributions subject to a prescribed divergence budget, and characterize the resulting tradeoff between embedding rate, decoding reliability, and statistical security. Experimentally, we compare our method against prior multi-bit watermarking schemes across modern language models and payload regimes. Our approach achieves substantially lower watermark detectability while maintaining competitive message-recovery performance and generation quality. Our results provide a unified view of secure multi-bit watermarking and recover several commonly used security notions as special cases.