🤖 AI Summary
This study investigates the impact of knowledge distillation (KD) on model memorization of fine-tuning data and its privacy implications. Focusing on the canonical setting of distilling large teacher models into smaller student models, we introduce the analytical lens of “memory transitivity” to systematically evaluate how diverse KD techniques—including logits-based, hidden-state-based, and attention-based methods—suppress student models’ memorization of task-specific fine-tuning data. Experimental results show that KD-trained student models exhibit an average 42% reduction in memorization rate while retaining over 98% of the teacher’s task performance. Crucially, this work provides the first empirical evidence that KD inherently confers implicit privacy protection: by transferring distilled knowledge rather than raw data patterns, it attenuates students’ fidelity to original fine-tuning examples. These findings establish a novel theoretical foundation and empirical support for lightweight, efficient, and privacy-enhanced model compression.
📝 Abstract
Large language models (LLMs) are known to memorize parts of their training data, raising important concerns around privacy and security. While previous research has focused on studying memorization in pre-trained models, much less is known about how knowledge distillation (KD) affects memorization.In this study, we explore how different KD methods influence the memorization of fine-tuned task data when a large teacher model is distilled into smaller student variants.This study demonstrates that distilling a larger teacher model, fine-tuned on a dataset, into a smaller variant not only lowers computational costs and model size but also significantly reduces the memorization risks compared to standard fine-tuning approaches.