🤖 AI Summary
Supervised fine-tuning of large language models (LLMs) exhibits memory bias, increasing risks of leakage of private or copyrighted information—posing critical security and privacy threats. Method: We first formally characterize the strong skewness of memory distribution and establish a theoretical linkage between memory probability and the generation process. We propose a memory-probability modeling framework grounded in sequence length and inter-sample similarity, enabling interpretable, quantitative, and decoupled memory-risk assessment. Further, we design a hierarchical detection strategy for early identification and mitigation. Results: Extensive experiments across multiple mainstream LLMs demonstrate that memory is highly concentrated in a tiny fraction of training samples; our proposed metrics are reproducible, significantly improving both detection accuracy and intervention timeliness for memorization leakage.
📝 Abstract
Memorization in Large Language Models (LLMs) poses privacy and security risks, as models may unintentionally reproduce sensitive or copyrighted data. Existing analyses focus on average-case scenarios, often neglecting the highly skewed distribution of memorization. This paper examines memorization in LLM supervised fine-tuning (SFT), exploring its relationships with training duration, dataset size, and inter-sample similarity. By analyzing memorization probabilities over sequence lengths, we link this skewness to the token generation process, offering insights for estimating memorization and comparing it to established metrics. Through theoretical analysis and empirical evaluation, we provide a comprehensive understanding of memorization behaviors and propose strategies to detect and mitigate risks, contributing to more privacy-preserving LLMs.