AuxMark: Defending Against Unauthorized Agent Distillation via Auxiliary Behavioral Watermarking

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unauthorized distillation of large language model (LLM) agent trajectories and the inadequacy of existing watermarking techniques in interactive environments. We propose a behavior-based watermarking framework tailored for LLM agents. By dynamically injecting auxiliary behaviors and storing private evidence cards, the method constructs paired probes for black-box auditing, enabling both model-level detection and trajectory attribution while preserving structured interaction properties and exhibiting robustness against data sanitization. Experimental results demonstrate that the proposed framework successfully identifies all 24 distilled models across multiple benchmarks with zero false positive rate, while maintaining lossless task utility.
📝 Abstract
Large language model agents can acquire complex capabilities through multi-step interaction and tool use, but their trajectories can also be illegally collected to dis- till student agents. However, existing watermarking methods either do not fit the structured and interactive nature of agent environments or lack reliable effective- ness across tasks and model architectures. We introduce AuxMark, a behavioral watermarking framework for tracing unauthorized agent distillation. AuxMark dynamically inserts safe, non-essential auxiliary action into teacher trajectories, and stores the associated contexts as private evidence cards. To audit a suspicious student model, AuxMark constructs paired real and fake probes from these cards and applies a card-level sign test. This black-box protocol supports both model- level detection and trace-level attribution. Across three agent benchmarks, two teacher agents, and four student architectures, AuxMark detects all 24 distilled models with zero false positives on 48 clean models. It also preserves task utility and remains effective against data flooding, paraphrasing, truncation, and adaptive cleaning attacks. Our code will be released at this URL.
Problem

Research questions and friction points this paper is trying to address.

agent distillation
behavioral watermarking
unauthorized distillation
large language model agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavioral Watermarking
Agent Distillation
Black-box Audit
Auxiliary Action
Trace Attribution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiqing Feng
Xidian University
H
Haozhe Feng
Zhejiang University
S
Shunan Shang
Xidian University
Xiaoyu Zhang
Xiaoyu Zhang
Xidian University
machine learningdata security
J
Jian Lou
Sun Yat-sen University
Haodong Zhao
Haodong Zhao
Shanghai Jiao Tong University
Federated LearningLLM
Mingxun Zhou
Mingxun Zhou
Hong Kong University of Science and Technology
Security and Privacy