LLM Watermark Evasion via Bias Inversion

📅 2025-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the insufficient robustness of large language model (LLM) watermarking techniques under adversarial settings, this paper proposes Bias-Inversion Rewriting Attack (BIRA), a theory-driven, model-agnostic watermark evasion method. BIRA requires no access to watermark algorithm details or model internals; instead, it analyzes output logits distributions to identify and suppress token-level biases exploitable by watermarking mechanisms, then applies semantics-preserving LLM rewriting to achieve effective evasion. Evaluated across multiple mainstream statistical watermarking schemes, BIRA achieves an average evasion rate exceeding 99%, while preserving semantic fidelity significantly better than baseline attacks. This work establishes the first general-purpose watermark evasion framework, systematically exposing structural vulnerabilities inherent in existing statistical watermarks. It provides critical theoretical insights and empirical benchmarks for designing watermarking schemes resilient to adversarial manipulation.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Ethics — Bias, Fairness, Transparency & Privacy

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Watermarking for large language models (LLMs) embeds a statistical signal during generation to enable detection of model-produced text. While watermarking has proven effective in benign settings, its robustness under adversarial evasion remains contested. To advance a rigorous understanding and evaluation of such vulnerabilities, we propose the emph{Bias-Inversion Rewriting Attack} (BIRA), which is theoretically motivated and model-agnostic. BIRA weakens the watermark signal by suppressing the logits of likely watermarked tokens during LLM-based rewriting, without any knowledge of the underlying watermarking scheme. Across recent watermarking methods, BIRA achieves over 99% evasion while preserving the semantic content of the original text. Beyond demonstrating an attack, our results reveal a systematic vulnerability, emphasizing the need for stress testing and robust defenses.
Problem

Research questions and friction points this paper is trying to address.

Proposes Bias-Inversion Rewriting Attack to evade LLM watermarks
Reduces watermark signal by suppressing likely watermarked tokens
Reveals systematic vulnerability in current watermarking methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bias-Inversion Rewriting Attack weakens watermark signal
Suppresses logits of likely watermarked tokens during rewriting
Achieves over 99% evasion while preserving semantics