🤖 AI Summary
This study addresses the vulnerability of large language model watermarks to localized tampering and their inability to pinpoint such alterations. We propose a watermarking framework that supports local integrity verification by incorporating error-correcting code constraints and boundary anchor mechanisms, coupled with a dynamic programming decoder. This approach extends conventional source attribution to fine-grained edit localization while enabling a configurable trade-off between detection reliability and generation quality. Experimental results demonstrate that under mixed-editing scenarios, the proposed method achieves a block-level true positive rate of 99.7% with a false positive rate below 7.6%. It effectively distinguishes original content from tampered segments while maintaining low perplexity, thereby offering a robust solution for verifying the provenance and integrity of watermarked text at a granular level.
📝 Abstract
LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified content is still attributed to the original model. We propose Anchor-ECC, which incorporates the error-correcting code (ECC) constraints and explicit boundary anchors into the watermark structure and pairs them with a dynamic-programming decoder to detect and localize post-generation edits. Across Qwen3-8B, Mistral-7B-Instruct-v0.3, and OPT-125M, the approximate-hard setting achieves about 99.7% block-level true positive rate (TPR) with at most 7.6% false alarm rate (FAR) for edit detection under mixed insertions, deletions, and substitutions, while preserving the distinction between watermarked outputs and unwatermarked text. Additional quality experiments identify lower-perplexity configurations that retain strong edit-detection performance. Together, these results extend LLM watermarking from source identification to local integrity verification while supporting configurable trade-offs between detection reliability and generation quality.