Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of quantifying information flow and precisely selecting training samples in self-distillation without external teachers. To this end, it proposes InFlow, a framework that formally conceptualizes self-distillation as a potential-to-realized information flow metric for the first time. By integrating confidence calibration trajectories, Jensen-Shannon divergence, and belief shifts, the method constructs a retrieval-guided mechanism for dynamic online sample selection. Extensive experiments demonstrate that InFlow achieves state-of-the-art cross-model average performance across four open-source models and three knowledge domains. Ablation studies further validate the effectiveness of its two-stage design. Overall, this work establishes a novel paradigm for self-improvement through self-distillation.
📝 Abstract
Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen--Shannon divergence between the teacher's initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher--student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at https://github.com/1240148048/INFLOW.
Problem

Research questions and friction points this paper is trying to address.

Self-Distillation
Information Flow
Training Example Selection
Knowledge Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Information Flow
Retrieval-Guided
On-Policy Distillation
Jensen-Shannon Divergence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rui Wang
Peking University
R
Ruijie Wang
University of Oxford
B
Bo Chen
The University of Hong Kong
J
Jiangxuan Long
The University of Hong Kong
Yingyu Liang
Yingyu Liang
The University of Hong Kong
machine learning