Score
Designs and implements metrics, measurement protocols, and experiments to quantify how models memorize training data, covering memorization amplification, per-sample memorization scores, and unintended memorization. Builds analyses and tools for outlier memorization assessment (for example using per-sample gradient norms), compares memorization under different training conditions, and correlates memorization measurements with risks such as privacy leakage.
Large language models (LLMs) exhibit “undesired memorization”—the excessive retention and leakage of sensitive training data fragments—posing critical risks for privacy violations and membership inference attacks. Method: We propose the first three-dimensional taxonomy (granularity, retrievability, extractability) to systematically characterize this phenomenon; establish a privacy–utility trade-off analysis framework unifying exposure, membership inference, and related metrics; and extend the study to emerging paradigms including retrieval-augmented generation (RAG) and diffusion language models. Through systematic literature review, empirical attribution analysis, and defense evaluation, we construct a structured knowledge graph and an open-source, dynamically updated literature repository. Results: Our work identifies six key frontiers in LLM memorization governance, advancing the field from ad hoc practice toward rigorous, systematized science.
This paper uncovers a fundamental tension between memorization and trustworthiness (i.e., fairness, robustness, and privacy) in machine learning, revealing that existing research conflates three distinct long-tail phenomena: inter-class imbalance, intra-class atypicality, and label noise—leading to misidentification and misguided suppression of memorization. To address this, we propose the first “three-level granularity” analytical framework for long-tail distributions, disentangling the functionally heterogeneous roles of memorization across granularities: it must be preserved to ensure fairness but suppressed to enhance robustness and privacy. Leveraging distributional modeling, theoretical analysis, and cross-domain synthesis, we establish a granularity-aware taxonomy and a principled trade-off evaluation paradigm. Our work redefines the theoretical foundations of trustworthy ML and delivers systematic design principles and a practical roadmap for context-sensitive memorization control.
This paper systematically investigates memorization phenomena in large language models (LLMs) and their associated privacy and ethical risks. Addressing the challenges of complex memorization mechanisms, difficulty in detection, and fragility of mitigation strategies, we propose an integrated “mechanism–detection–mitigation” framework. Our approach enables fine-grained memorization localization via prefix extraction and membership inference attacks; combines differential privacy during training with post-training model unlearning to jointly optimize utility and privacy; and—firstly—formally defines and empirically characterizes the boundary between learning and memorization. Extensive experiments across mainstream LLMs and benchmark datasets reveal dynamic memorization evolution across repeated pretraining and fine-tuning stages. Results demonstrate that our method significantly reduces membership inference success rates by 37.2% on average while preserving model performance, establishing a reproducible, scalable technical pathway and theoretical benchmark for trustworthy LLM development.
Supervised fine-tuning of large language models (LLMs) exhibits memory bias, increasing risks of leakage of private or copyrighted information—posing critical security and privacy threats. Method: We first formally characterize the strong skewness of memory distribution and establish a theoretical linkage between memory probability and the generation process. We propose a memory-probability modeling framework grounded in sequence length and inter-sample similarity, enabling interpretable, quantitative, and decoupled memory-risk assessment. Further, we design a hierarchical detection strategy for early identification and mitigation. Results: Extensive experiments across multiple mainstream LLMs demonstrate that memory is highly concentrated in a tiny fraction of training samples; our proposed metrics are reproducible, significantly improving both detection accuracy and intervention timeliness for memorization leakage.
This study investigates the impact of knowledge distillation (KD) on model memorization of fine-tuning data and its privacy implications. Focusing on the canonical setting of distilling large teacher models into smaller student models, we introduce the analytical lens of “memory transitivity” to systematically evaluate how diverse KD techniques—including logits-based, hidden-state-based, and attention-based methods—suppress student models’ memorization of task-specific fine-tuning data. Experimental results show that KD-trained student models exhibit an average 42% reduction in memorization rate while retaining over 98% of the teacher’s task performance. Crucially, this work provides the first empirical evidence that KD inherently confers implicit privacy protection: by transferring distilled knowledge rather than raw data patterns, it attenuates students’ fidelity to original fine-tuning examples. These findings establish a novel theoretical foundation and empirical support for lightweight, efficient, and privacy-enhanced model compression.
This work addresses the challenge of distinguishing genuine memorization from statistical generalization in large language models, a distinction that existing methods often conflate, leading to overestimation of training data leakage risks. The authors propose a lightweight, retraining-free criterion grounded in a novel theoretical framework that integrates model priors. By comparing the generation probabilities of candidate suffixes under a specific prefix versus irrelevant prompts, their approach effectively disentangles memorization from generalization. Experimental evaluations on models such as LLaMA and OPT—leveraging probability contrast, data tracing, and prior modeling—reveal that 55%–90% of sequences previously misclassified as memorized are in fact common statistical patterns. Notably, even approximately 40% of sequences appearing only once are attributable to generalization, demonstrating that frequency alone is an unreliable indicator of memorization and substantially improving both the accuracy and efficiency of leakage assessment.
该文指出通过微调记忆化结果来解决版权问题的方法存在测量程序无效、缺乏对照实验等问题,其声称的从训练数据中提取大量版权书籍内容的说法未得到支持。
Conventional correlational analyses fail to establish causal links between training data and language model (LM) behavior. Method: We propose a “rewriting history” intervention framework that systematically identifies, modifies, and re-trains on training documents containing target knowledge—leveraging co-occurrence statistics and information retrieval for precise document matching—and quantifies behavioral changes via standardized benchmarks. Contribution/Results: This work introduces the first controlled, causal data intervention at the training stage, moving beyond observational studies to enable rigorous causal testing of data effects on LM behavior. Experiments demonstrate that localized data rewriting significantly alters model knowledge expression; however, current matching strategies remain insufficient to fully account for knowledge acquisition, revealing the inherent complexity of the mapping between training data and emergent model knowledge.
Large language models (LLMs) in federated learning (FL) pose cross-client training data memorization risks; existing detection methods focus solely on single-sample memorization, neglecting fine-grained inter-sample memorization, and centralized evaluation techniques do not directly transfer to FL. Method: We extend fine-grained cross-sample memorization assessment to FL for the first time, proposing a unified analytical framework that quantifies both intra-client and cross-client memorization. We systematically investigate the impact of decoding strategies, prefix length, training rounds, and FL algorithms on memorization behavior. Results: Experiments confirm that FL-trained LLMs indeed memorize client-specific data, with intra-client memorization significantly stronger than cross-client memorization. Key training and inference factors exert quantifiable, non-negligible effects on memorization intensity. This work establishes a novel, empirically grounded methodology for privacy risk assessment in FL, enabling principled evaluation of model memorization across heterogeneous clients.
The boundary between memorization (reproducing training data) and generalization (generating novel samples) in diffusion models remains poorly understood, yet this boundary critically determines copyright and privacy risks. Method: We establish a theoretical framework that—first for underparameterized diffusion models—reveals the critical mechanism governing the memorization–generalization phase transition: the transition is governed by the relative weighting of memory and generalization components in the training loss. We derive an analytical criterion to precisely predict the critical model size at which memorization dominates, and design a “mathematical laboratory” using synthetic and structured images to quantify the evolution of loss component weights via both theoretical analysis and gradient descent experiments. Contribution/Results: Experiments validate the high accuracy of our theoretical predictions, yielding explicit model capacity thresholds beyond which memorization becomes likely. This work provides the first analytically tractable and empirically verifiable theoretical foundation for designing safe, controllable diffusion models.
This work addresses a critical gap in evaluating memorization in large language models (LLMs), as existing approaches predominantly focus on forced leakage of training data while overlooking the models’ propensity to leak such data during ordinary usage. To bridge this gap, the authors propose PropMe, a novel “propensity-aware” memorization evaluation framework that integrates both adversarial prefix attacks and non-adversarial scenarios. PropMe employs SimpleTrace—a lightweight tracing pipeline built upon Infinigram—to quantify verbatim, approximate, and propensity-based memorization. A new metric, propensity conversion, is introduced to disentangle a model’s inherent memorization capacity from its tendency to leak memorized content. Experimental results demonstrate that while LLMs readily leak data under adversarial prompts, their leakage propensity remains remarkably low in standard usage; furthermore, continued pretraining effectively reduces both memorization capacity and leakage propensity.
This work addresses the opacity of training data in large language models for code, which complicates the assessment of data leakage risks. The authors propose a perturbation-based quantitative method to systematically measure a model’s “memorization advantage”—defined as the performance gap between inputs the model may have encountered during training and those it has not. By integrating perturbation testing, cross-model comparisons, and a multi-task benchmark encompassing code generation, comprehension, vulnerability detection, and repair, the study reveals that memorization advantage varies significantly with task type and model architecture. Experiments show that StarCoder exhibits notable memorization advantage on certain tasks, whereas QwenCoder demonstrates substantially less. Moreover, widely used datasets such as CVEFixes and Defects4J exert minimal memorization effects, suggesting that models primarily rely on generalization rather than rote memorization.