🤖 AI Summary
This study addresses the vulnerability of conventional watermarking methods in masked diffusion models to frequency attacks and their reliance on sequential generation order. To overcome these limitations, this work proposes TANGO, a watermarking framework that embeds color biases into newly generated tokens based on already revealed neighboring tokens. By employing a dynamic coloring mechanism rather than a fixed greenlist, TANGO effectively disrupts positional correlations. Furthermore, it leverages a secret key to partition vocabulary color classes in conjunction with local context, enabling robust watermarking for parallel, unordered generation. Experimental results demonstrate that TANGO achieves high detection rates across two distinct model architectures while effectively resisting frequency analysis attacks. Crucially, the method preserves the statistical properties of the generated text, maintaining distributions closely aligned with those of unwatermarked outputs.
📝 Abstract
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.