🤖 AI Summary
This study addresses a critical side effect of production-grade text watermarking for large language models: while preserving detectability, watermarking induces factual hallucinations, causing models to generate incorrect content that disregards evidence. This work systematically quantifies this phenomenon for the first time, revealing that it stems from a dual failure mechanism involving token perturbation and attention drift. To mitigate this issue, we propose a plug-and-play correction strategy based on token reweighting and attention optimization, which is compatible with mainstream watermarking algorithms such as RAG and KGW. Experimental results demonstrate that the proposed method reduces factual errors by approximately 90% while maintaining both watermark detection rates and text fluency. Furthermore, this work establishes factuality as a core evaluation metric for text watermarking systems.
📝 Abstract
Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled retrieval-augmented generation setting, we compare unwatermarked and watermarked generations under the same context, query, and decoding configuration, and quantify their factual accuracy decrease. Across six representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID watermarking, we consistently observe watermark-induced hallucination. Watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) token perturbations in the current-step arising from method-specific reweighting or keyed sampling, and (2) prefix-induced attention drift, which accumulates through autoregressive decoding and weakens later attention to the factual context. Motivated by this analysis, we propose two plug-in interventions at the token and attention levels that can be integrated into existing watermarking methods to improve factuality. At a matched TPR of 0.90 at 1% FPR, combining the two interventions reduces factual errors by approximately 90% relative to watermark-only decoding while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.