🤖 AI Summary
This study addresses the challenge that large language model (LLM) watermark detectors are rarely made public due to attack concerns, thereby limiting technical transparency and ecosystem interoperability. To overcome this, we propose the first split-key watermarking mechanism that decouples public detection from private verification. By releasing a partial key for transparent detection while retaining a private key for final verification, and by introducing a statistical test based on score imbalance to precisely identify tampering, our approach ensures robust security. Theoretical analysis and empirical results demonstrate that publicly releasing the half-key detector does not significantly increase removal attack risks while effectively defending against forgery. This work achieves transparent watermark detection without compromising security, lowers liability barriers for service providers, and establishes a new paradigm for the open deployment of LLM watermarking technologies.
📝 Abstract
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between public and private scores. We introduce a statistical test for this imbalance, and combine it with the full key verdict in a two-stage mechanism. Third, we evaluate the split-key method on a wide range of removal and forgery attacks, comparing the uninformed to detector-informed settings. Public detection improves removal only at small edit budgets, since plain rephrasing already strips the watermark at a lower quality cost, but it does enable forgery, which the private pipeline can identify. Overall, releasing half of the watermark enables transparency and interoperability, and tampering with the released half stays detectable. This bounds the provider's liability and questions the need to keep detectors fully private.