Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate
This study addresses a critical side effect of production-grade text watermarking for large language models: while preserving detectability, watermarking induces factual hallucinations, causing models to generate incorrect content that disregards evidence. This work systematically quantifies this phenomenon for the first time, revealing that it stems from a dual failure mechanism involving token perturbation and attention drift. To mitigate this issue, we propose a plug-and-play correction strategy based on token reweighting and attention optimization, which is compatible with mainstream watermarking algorithms such as RAG and KGW. Experimental results demonstrate that the proposed method reduces factual errors by approximately 90% while maintaining both watermark detection rates and text fluency. Furthermore, this work establishes factuality as a core evaluation metric for text watermarking systems.