RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language model (LLM) watermarks to deletion attacks, where positional shifts compromise detection robustness. To overcome this limitation, we propose a watermarking method based on Reed–Muller codes. Departing from conventional global recovery paradigms, our approach injects local algebraic structures through secret-key vocabulary partitioning and achieves robust detection by exploiting local Reed–Solomon consistency induced by affine line restrictions. Efficient verification is realized by integrating the Berlekamp–Welch algorithm with subsequence low-degree testing. Experimental evaluations on the C4 and ELI5 datasets demonstrate that the proposed method maintains high detection rates under diverse deletion and rewriting attacks, significantly outperforming existing baselines.
📝 Abstract
Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at https://github.com/BaichengDanny/RMCW.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model watermarking
deletion attacks
robustness
post-processing attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reed-Muller codes
deletion-robust watermarking
large language models
local algebraic structure
Berlekamp-Welch test
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yi Wang
Institute for Interdisciplinary Information Sciences, Tsinghua University; University of Wisconsin–Madison
Baicheng Chen
Baicheng Chen
University of California San Diego
MetasurfaceMetamaterialWireless SensingMobile HealthSecurity/Privacy
Y
Yu Wang
Institute of Information Engineering, Chinese Academy of Sciences
J
Jian Zhao
Institute for Interdisciplinary Information Sciences, Tsinghua University; Xiongan AI Institute
Y
Yilei Chen
Institute for Interdisciplinary Information Sciences, Tsinghua University; Shanghai Qi Zhi Institute
Tianxing He
Tianxing He
Tsinghua University
NLP