When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that pointwise forward KL clipping in online self-distillation suppresses error-correction signals and reinforces repetitive patterns, causing models to deviate from the teacher distribution. Building upon an online self-distillation framework with forward KL divergence optimization, this work combines controlled comparative experiments with theoretical analysis of probability distributions to provide the first mathematical proof that such clipping strategies can reverse the direction of correction under specific conditions. Both theoretical and empirical findings confirm that pointwise KL clipping induces persistent repetitive generation and causes internal probability allocations to diverge from the teacher model. By revealing the fundamental failure mechanism underlying this approach, this research offers a critical cautionary insight for optimizing knowledge distillation objectives.
📝 Abstract
On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clipping of the forward KL objective stabilizes training. Pointwise clipping caps each vocabulary-wise forward KL term at a fixed threshold before summing over the vocabulary. Follow-up studies have adopted this clipping, but its effect on training has not been directly examined. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to the end of the response than their unclipped counterparts. We trace this failure to the clipped objective. We prove that the clipped objective can fail to correct the student toward the teacher and can instead push clipped and unclipped token probabilities away from its teacher. Our training runs agree with this analysis: inside repetitions, the clipped student places less probability than its teacher on leaving the repetition, and more on continuing it, whereas the unclipped runs stay close to their teachers.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
pointwise clipping
forward KL divergence
repetition failure
mathematical reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy self-distillation
Pointwise clipping
Forward KL divergence
Repetition failure
Mathematical reasoning
🔎 Similar Papers
No similar papers found.