🤖 AI Summary
This study addresses the vulnerability of crowdsourced fine-tuning for large language models, where malicious actors can covertly amplify privacy leakage risks using minimal poisoned data that evades existing defenses. By investigating how topic-based poisoning enhances training data extraction under black-box access, the authors conduct a systematic analysis employing supervised fine-tuning, adversarial data poisoning, and multiple filtering evaluation techniques. The work demonstrates for the first time that merely 50 poisoned samples can increase data extraction rates by over threefold in Qwen and Llama models. Furthermore, the best-performing filter achieves an F1 score of only 0.378, underscoring both the stealthiness of this attack and the significant limitations of current defensive mechanisms against targeted privacy exploitation in collaborative model training scenarios.
📝 Abstract
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches $3.71\times$ the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and $3.08\times$ for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.