Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of teacher supervision in self-distillation, where credit direction and magnitude are often coupled and highly sensitive to judgment errors. We propose the first decoupled-credit self-distillation method that determines credit direction via belief marginal probing and quantifies contribution magnitude through marginal information gain, enabling precise step-to-token credit assignment. Integrated with reinforcement learning verification rewards, online preference self-distillation, and policy optimization algorithms, the proposed approach outperforms mainstream baselines across 11 benchmarks. Notably, it achieves improvements of 8.45 and 7.01 points on mathematical reasoning and multimodal tasks, respectively, while successfully correcting the credit direction for 6% of tokens.
📝 Abstract
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.
Problem

Research questions and friction points this paper is trying to address.

Self-Distillation
Credit Assignment
Reinforcement Learning
Teacher Supervision
Policy Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupled Credit Self-Distillation
Belief-Margin Probing
Marginal Information Gain
Token-level Credit Assignment
Policy Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yugu Li
School of CSIT, Adelaide University, Adelaide, SA 5000, Australia
Z
Zehong Cao
School of CSIT, Adelaide University, Adelaide, SA 5000, Australia
Peizhen Li
Peizhen Li
PhD candidate, Macquarie University
RoboticsMachine LearningData Mining
Y
Yang Zhang
CAIAA, University of North Texas, Denton, TX 76203, USA
Siyi Hu
Siyi Hu
Adelaide University
Generative AIReinforcement LearningMulti-Agent Systems
Jianglin Qiao
Jianglin Qiao
University of South Australia
Artifical Intelligence