Language Model Fingerprinting Requires Rethinking Watermark Teachers

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between text quality and detection rates in large language model fingerprinting caused by watermarked teacher models. To this end, we propose an optimization framework based on statistical watermark distillation. Methodologically, leveraging the properties of aggregated query signals, we introduce a near-tie constraint mechanism based on logit differences to optimize signal placement. Furthermore, by integrating token surprisal analysis with top-1 relative logit difference constraints, our approach overcomes the limitations of traditional single-sample verification. Experimental results demonstrate that sparse signal aggregation significantly improves the Pareto frontier of detection performance and text quality. Moreover, the proposed framework effectively enhances the robustness of existing watermarking schemes across diverse models and deployment variations.
📝 Abstract
LLM fingerprinting via watermark distillation embeds a statistical watermark signal into model weights, enabling model owners to identify their models behind black-box APIs. Revisiting a recent protocol, we find that its utility evaluation understates text quality degradation in open-ended generation, favoring overly strong watermark teachers. Weakening the watermark improves text quality but sacrifices detectability. To move beyond this trade-off, we rethink whether text watermarks designed for verifying generated text are suitable distillation teachers for model fingerprinting. Such watermarks are typically designed to remain detectable from an individual output, limiting how sparse the watermark signal can be. In contrast, fingerprint verification can aggregate signal across queries, making sparser watermark signals viable. This raises a key question: where should the sparse signal be placed? We analyze signal placement through token surprisal and show that, even at comparable watermark strength, different placements can target tokens with different plausibility under the base model. This motivates near-tie restriction, which uses top-1-relative logit gaps to restrict the watermark bias to tokens close to the base model's top prediction. Across multiple models, near-tie improves detection--quality frontiers under deployment changes, preserves higher text quality across query budgets, and further improves existing watermarking schemes when combined with them.
Problem

Research questions and friction points this paper is trying to address.

LLM fingerprinting
watermark distillation
text quality
detectability
signal placement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Watermark Distillation
LLM Fingerprinting
Near-tie Restriction
Token Surprisal
Sparse Watermark Signal
🔎 Similar Papers