Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the differential effects of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) teachers on Online Policy Distillation (OPD) within the context of model merging. Through controlled single-teacher experiments, we compare the instructional efficacy of both teacher types across agent, reasoning, and perception tasks. Our findings reveal that RL teachers are significantly more effective than SFT teachers, primarily because their parameter space is closer to that of the student model, thereby facilitating easier alignment. Specifically, students guided by RL outperform those guided by SFT by 4.27, 1.50, and 0.86 percentage points on agent, reasoning, and perception tasks, respectively. This work highlights the critical role of parameter-space proximity in knowledge transfer efficiency and provides empirical guidance for teacher selection within OPD frameworks.
📝 Abstract
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinforcement learning (RL). Yet equally strong teachers need not be equally good teachers. We probe this choice through controlled single-teacher OPD, a building block of multi-teacher OPD: in Agentic, Reasoning, and Perception, comparably performing SFT and RL teachers are trained from Qwen3.5-9B, each guiding a student initialized from it. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points in Agentic, Reasoning, and Perception, respectively, and recover more of their teachers'performance gains over the base model. The contrast is clearest in Agentic, where the best SFT-guided student recovers only 44.44% of its teacher's gain, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our analysis points to an explanation: RL teachers stay much closer to the shared initialization in parameter space than SFT teachers and are therefore easier for their students to follow.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Model Merging
Supervised Fine-Tuning
Reinforcement Learning
Teacher-Student
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Model Merging
Reinforcement Learning
Supervised Fine-Tuning
Parameter Space
🔎 Similar Papers
No similar papers found.