PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement

📅 2025-08-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Commercial deployment of large language models (LLMs) faces challenges in copyright protection, as black-box watermarking schemes are vulnerable to removal via incremental fine-tuning and compromised by feature-space adversarial attacks. Method: This paper proposes PREE (Prefix-Enhanced Embedding for Robust Enforcement), a fingerprint editing framework that implicitly encodes copyright information into model weight perturbations via dual-channel knowledge editing and parameter-offset encoding, augmented by a prefix-based trigger mechanism to enhance robustness. Contribution/Results: Unlike existing high-complexity trigger-based methods, PREE achieves benign, adaptive, and low-overhead fingerprint embedding: it attains >90% trigger accuracy on LLaMA-3 and Qwen-2.5, modifies <0.03% of parameters, yields zero false positives, and maintains strong resilience against fine-tuning, pruning, and quantization. PREE significantly improves watermark stealth, security, and practical deployability.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Addressing the intellectual property protection challenges in commercial deployment of large language models (LLMs), existing black-box fingerprinting techniques face dual challenges from incremental fine-tuning erasure and feature-space defense due to their reliance on overfitting high-perplexity trigger patterns. Recent work has revealed that model editing in the fingerprinting domain offers distinct advantages, including significantly lower false positive rates, enhanced harmlessness, and superior robustness. Building on this foundation, this paper innovatively proposes a $ extbf{Pr}$efix-$ extbf{e}$nhanced Fingerprint $ extbf{E}$diting Framework (PREE), which encodes copyright information into parameter offsets through dual-channel knowledge edit to achieve covert embedding of fingerprint features. Experimental results demonstrate that the proposed solution achieves the 90% trigger precision in mainstream architectures including LLaMA-3 and Qwen-2.5. The minimal parameter offset (change rate < 0.03) effectively preserves original knowledge representation while demonstrating strong robustness against incremental fine-tuning and multi-dimensional defense strategies, maintaining zero false positive rate throughout evaluations.
Problem

Research questions and friction points this paper is trying to address.

Protecting intellectual property in large language models
Overcoming fingerprint erasure from incremental fine-tuning
Enhancing robustness against feature-space defense strategies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prefix-enhanced fingerprint editing framework
Dual-channel knowledge edit encoding
Minimal parameter offset preserving representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.