🤖 AI Summary
This study addresses the trade-off between message recovery and text quality in multi-bit watermarking, as well as the lack of decoding guarantees with bounded error rates. To this end, it proposes the CertMark framework, which introduces a distribution-preserving multi-bit watermarking mechanism based on Gumbel-max sampling to maintain the original distribution without distortion. Furthermore, it incorporates a certified abstention decoding rule with mathematically proven error bounds, integrating model-agnostic or model-aware decoders with probabilistic statistical certification techniques. Experimental results demonstrate that CertMark reliably recovers multi-bit messages while preserving the perplexity of unwatermarked text, achieving bit accuracy that significantly surpasses existing baseline methods.
📝 Abstract
Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accumulated token-level evidence, without a certified abstention rule that bounds the probability of outputting an incorrect message. We introduce CertMark, a distribution-preserving multi-bit watermark with certified decoding. Rather than modifying probabilities, CertMark uses the embedded message to seed an exact Gumbel-max sampler, thereby preserving the model's original sampling distribution. We propose two scalable decoders: a model-agnostic, text-only decoder and a model-aware variant that leverages the original next-token distributions for stronger recovery. Both support certified abstention with mathematical bounds on the probability of returning an incorrect message. Across text completion, summarization, and story generation, CertMark matches the perplexity of unwatermarked text while reliably recovering multi-bit messages. The model-aware decoder further achieves higher bit accuracy than probability-biasing baselines. Our code is publicly available at https://github.com/Batorskq/CertMark.