🤖 AI Summary
This study addresses the problem of language models memorizing and reproducing sensitive data during fine-tuning. It proposes Target Reference Advantage (TRA), a novel memorization metric, alongside the TRAP defense method. Specifically, this work pioneers the use of complementary reference models to define TRA, enabling the precise identification of rare sensitive text fragments. Furthermore, it introduces a one-sided differentiable regularization mechanism that imposes targeted penalties, combined with early stopping strategies and token-level analysis to mitigate privacy leakage. Experiments on essay and clinical datasets demonstrate that TRAP reduces model memorization rates to pre-training levels while preserving task utility almost entirely. These results indicate that the proposed approach significantly outperforms general regularization techniques and differential privacy baselines in preventing the extraction of sensitive information.
📝 Abstract
Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.