🤖 AI Summary
To address the insufficient robustness of large language model (LLM) watermarking techniques under adversarial settings, this paper proposes Bias-Inversion Rewriting Attack (BIRA), a theory-driven, model-agnostic watermark evasion method. BIRA requires no access to watermark algorithm details or model internals; instead, it analyzes output logits distributions to identify and suppress token-level biases exploitable by watermarking mechanisms, then applies semantics-preserving LLM rewriting to achieve effective evasion. Evaluated across multiple mainstream statistical watermarking schemes, BIRA achieves an average evasion rate exceeding 99%, while preserving semantic fidelity significantly better than baseline attacks. This work establishes the first general-purpose watermark evasion framework, systematically exposing structural vulnerabilities inherent in existing statistical watermarks. It provides critical theoretical insights and empirical benchmarks for designing watermarking schemes resilient to adversarial manipulation.
📝 Abstract
Watermarking for large language models (LLMs) embeds a statistical signal during generation to enable detection of model-produced text. While watermarking has proven effective in benign settings, its robustness under adversarial evasion remains contested. To advance a rigorous understanding and evaluation of such vulnerabilities, we propose the emph{Bias-Inversion Rewriting Attack} (BIRA), which is theoretically motivated and model-agnostic. BIRA weakens the watermark signal by suppressing the logits of likely watermarked tokens during LLM-based rewriting, without any knowledge of the underlying watermarking scheme. Across recent watermarking methods, BIRA achieves over 99% evasion while preserving the semantic content of the original text. Beyond demonstrating an attack, our results reveal a systematic vulnerability, emphasizing the need for stress testing and robust defenses.