Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the vulnerability of existing behavioral watermarks for LLM agents to tool renaming, observation paraphrasing, and forgery attacks. To overcome these limitations, we propose Semantic Behavioral Watermarking (SBW), which embeds owner identifiers through semantic action clustering and history-conditioned encoding. Furthermore, SBW introduces a bucketing mechanism with key-collision resistance to ensure the unpredictability of new bucket assignments. This work is the first to simultaneously resolve both watermark removal and forgery threats. Experimental results on ToolBench and ALFWorld demonstrate that SBW significantly improves detection rates against paraphrasing attacks while reducing adaptive forgery success rates to the false-positive baseline. The source code has been made publicly available.
๐Ÿ“ Abstract
Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanoviฤ‡ et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.
Problem

Research questions and friction points this paper is trying to address.

behavioral watermarking
LLM agents
paraphrase robustness
forgery resistance
provenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Behavioral Watermarking
LLM Agents
Paraphrase Robustness
Forgery Resistance
Keyed Collision-Resistant Binning
๐Ÿ”Ž Similar Papers
S
Suxin Ji
University of Pennsylvania
H
Hungtao Wan
Independent Researcher
S
Shaoxuan Chen
University of Massachusetts Amherst
An Zhang
An Zhang
University of Science and Technology
Generative ModelsTrustworthy AIAgentic AIRecommender System