🤖 AI Summary
This study addresses the calibration degradation in student models during large language model knowledge distillation, which is caused by teacher hallucinations and high-entropy predictions. To mitigate this, we propose CaRE-KD, a framework that replaces static distillation objectives with confidence-gated, uncertainty-adaptive optimization. Its core innovation lies in a dual-granularity design: at the token level, it dynamically blends and adaptively switches between forward and reverse KL divergences; at the batch level, it introduces a Revival mechanism to suppress unreliable supervisory signals, integrated with gradient analysis and selective distillation techniques. Extensive evaluations across eleven benchmarks demonstrate that CaRE-KD significantly outperforms existing baselines, yielding a 3.2% improvement in ROUGE-L for instruction following and a 1.7% accuracy gain on GSM8K.
📝 Abstract
Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $α$--$β$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.