Anchor-ECC: Local Integrity Checking for Watermarked LLM Outputs via Error-Correcting Codes

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language model watermarks to localized tampering and their inability to pinpoint such alterations. We propose a watermarking framework that supports local integrity verification by incorporating error-correcting code constraints and boundary anchor mechanisms, coupled with a dynamic programming decoder. This approach extends conventional source attribution to fine-grained edit localization while enabling a configurable trade-off between detection reliability and generation quality. Experimental results demonstrate that under mixed-editing scenarios, the proposed method achieves a block-level true positive rate of 99.7% with a false positive rate below 7.6%. It effectively distinguishes original content from tampered segments while maintaining low perplexity, thereby offering a robust solution for verifying the provenance and integrity of watermarked text at a granular level.
📝 Abstract
LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified content is still attributed to the original model. We propose Anchor-ECC, which incorporates the error-correcting code (ECC) constraints and explicit boundary anchors into the watermark structure and pairs them with a dynamic-programming decoder to detect and localize post-generation edits. Across Qwen3-8B, Mistral-7B-Instruct-v0.3, and OPT-125M, the approximate-hard setting achieves about 99.7% block-level true positive rate (TPR) with at most 7.6% false alarm rate (FAR) for edit detection under mixed insertions, deletions, and substitutions, while preserving the distinction between watermarked outputs and unwatermarked text. Additional quality experiments identify lower-perplexity configurations that retain strong edit-detection performance. Together, these results extend LLM watermarking from source identification to local integrity verification while supporting configurable trade-offs between detection reliability and generation quality.
Problem

Research questions and friction points this paper is trying to address.

LLM watermarking
local integrity checking
post-generation edit detection
edit localization
source attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM watermarking
error-correcting codes
local integrity checking
dynamic-programming decoder
edit localization
🔎 Similar Papers
2023-10-27IACR Cryptology ePrint ArchiveCitations: 38