π€ AI Summary
This study addresses the limitations of existing semantic watermarks for large language models, including semantic narrowing due to preference misalignment, high resampling overhead, and degraded generation quality. We propose HammingMark, which introduces a novel dynamic verification mechanism over a compact hash space. Specifically, it constructs dynamic centers using preceding-sentence hashes and defines watermark validity via Hamming neighborhoods, enabling coarse-grained many-to-one mappings that accommodate diverse semantic expressions. Furthermore, semantic hashing combined with distance-constrained algorithms efficiently filters candidate tokens. Evaluated on C4 and BookSum, HammingMark achieves high robustness and detection rates while requiring only 2.2 sampling attempts per sentence, improving efficiency by 72.8% and preserving near-watermark-free generation quality.
π Abstract
Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.