KDFP: A first-principles approach to knowledge distillation in large language models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the research gap in general knowledge distillation for large language models and the demand for efficient private deployment on edge devices. Grounded in first principles, we propose KDFP, a white-box distillation framework that pioneers a paradigm for transferring general knowledge. The method constructs an optimization strategy by dynamically evaluating prior experience alongside real-time exploration, and introduces an instantaneous parameter reduction technique to substantially lower computational overhead. Experimental results demonstrate that KDFP achieves performance improvements of 1.6%–4.9% across nine benchmarks while increasing training efficiency by up to 99.1%, realizing a significant synergistic optimization between distillation accuracy and computational efficiency.
📝 Abstract
Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% $-$ 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
Large Language Models
General Knowledge
Edge Devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
Large Language Models
White-box Distillation
Ephemeral Parameter Reduction
First-principles Approach
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryan Swift
Thomas Lord Department of Computer Science, University of Southern California, Los Angeles, CA 90007, USA
Konstantinos Psounis
Konstantinos Psounis
Professor of Electrical and Computer Engineering and Computer Science, University of Southern
PerformancePrivacyMachine LearningDistributed Systems