Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
This study investigates the differential effects of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) teachers on Online Policy Distillation (OPD) within the context of model merging. Through controlled single-teacher experiments, we compare the instructional efficacy of both teacher types across agent, reasoning, and perception tasks. Our findings reveal that RL teachers are significantly more effective than SFT teachers, primarily because their parameter space is closer to that of the student model, thereby facilitating easier alignment. Specifically, students guided by RL outperform those guided by SFT by 4.27, 1.50, and 0.86 percentage points on agent, reasoning, and perception tasks, respectively. This work highlights the critical role of parameter-space proximity in knowledge transfer efficiency and provides empirical guidance for teacher selection within OPD frameworks.