Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization instability and performance collapse issues inherent in feedback-driven self-distillation for large language models by proposing FIRE, a dual-branch framework. To our knowledge, this work is the first to introduce softmax Fisher trajectories for computing token-level radii, effectively decoupling update directions from step sizes. Furthermore, FIRE mitigates interference from detrimental supervisory signals through a combination of online policy self-distillation, supervised fine-tuning (SFT) reweighting, and Fisher information-based feedback recalibration. Experimental results demonstrate that FIRE significantly enhances training stability, maintaining robust downstream task performance even in scenarios where standard methods fail.
📝 Abstract
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
Problem

Research questions and friction points this paper is trying to address.

self-distillation
large language models
optimization instability
performance collapse
feedback-based learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Fisher Information
Large Language Models
Feedback-Based Learning
Recalibration
🔎 Similar Papers
No similar papers found.