Forging LLM Authorship Fingerprints with Targeted Rewriting

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of LLM text attribution classifiers to adversarial rewriting attacks by demonstrating that fingerprint detectability does not guarantee provenance authenticity. We propose ForgePrint, a framework introducing a novel "search-distillation" paradigm for targeted fingerprint forgery. Specifically, it first transfers model output fingerprints to a target model via adversarial search, then employs knowledge distillation to train a 4B-parameter student model capable of single-pass conditional generation. Evaluated on the CNN/DM summarization task, our approach achieves a targeted success rate of 70.2%, substantially outperforming both the teacher model and existing baselines. Furthermore, we demonstrate successful fingerprint transfer to commercial APIs. These findings establish that purely textual attribution evidence remains unreliable under adversarial conditions, highlighting critical security implications for current LLM fingerprinting systems.
📝 Abstract
Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this problem as targeted fingerprint transfer: rewriting one model's output so that attribution classifiers assign it to a chosen target model. We study summarization, where different models receive the same document and express the same underlying content, providing a controlled setting for conditional generation. We introduce ForgePrint, a search-then-distil framework that first searches for rewrites that move attribution toward a target fingerprint, then distils the selected rewrites into a one-pass 4B Student model. On CNN/DM, the Student reaches 70.2% target success rate, outperforming both its Teacher (54.1%) and the strongest of six published rewriting baselines (39.3%), against held-out classifiers that are never queried by the attack. It also reaches 68.3% target success when transferring summaries from an open model toward chosen commercial models. These results show that fingerprint detectability should not be conflated with source authenticity, and that text-only attribution can provide misleading evidence of model identity under targeted rewriting, even when it is accurate on unmodified text.
Problem

Research questions and friction points this paper is trying to address.

model attribution
targeted rewriting
authorship fingerprint
provenance
LLM
Innovation

Methods, ideas, or system contributions that make the work stand out.

Targeted Fingerprint Transfer
Search-then-Distil Framework
Model Attribution Evasion
Knowledge Distillation
LLM Authorship Forgery