🤖 AI Summary
This study addresses the risk that fine-tuning may reactivate forgotten privacy vulnerabilities in large language models (LLMs), noting that existing attack methods rely on authentic private data that is difficult to obtain. To overcome this limitation, this work proposes ReGap, a data-free attack framework that reproduces privacy leakage without requiring genuine private supervision or target answers, relying solely on task structures and LLM-generated candidate signals. Specifically, the method leverages LLM generation techniques, an answer-token likelihood filtering mechanism, and Low-Rank Adaptation (LoRA) to recover private associations concealed within the model. Experimental results demonstrate that ReGap improves target association recovery rates by 6 to 21 percentage points across multiple models. These findings confirm that prior exposure significantly influences post-fine-tuning recoverability, thereby revealing latent privacy risks inherent in conventional customized fine-tuning practices.
📝 Abstract
Beyond adapting Large Language Models (LLMs) to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same distribution, i.e., the previous training dataset. We argue that such recovery remains possible without such impractical knowledge. We show that LLM-generated candidates can provide sufficient supervision to recover previously learned private associations. Based on this, we propose ReGap, a data-free attack that recovers private associations using task structure, filters them by answer-token likelihood, and updates the target model via low-rank adaptation. Specifically, ReGap requires neither target answers nor auxiliary genuine private supervision. Across six GPT-2, OPT, and Qwen3 models, ReGap improves target-association recovery by 6-21 percentage points over the post-training target model. Recovery remains substantial even when the adaptation identities are disjoint from all memorized and evaluation identities, with no exact target answers appearing in the generated or selected supervision. Moreover, the same trained adapters increase recovery from 42\% to 63\% on a previously exposed checkpoint, but produce no gain on a matched checkpoint that never encountered the targets. This contrast shows that adaptation alone is insufficient to explain the observed recovery and that prior target exposure strongly affects post-adaptation recoverability. Our findings highlight that routine model customization can reawaken latent privacy risks, warranting urgent attention from the academic and industrial communities.