🤖 AI Summary
This study investigates the impact of prompt data on teacher-student knowledge transfer in online policy distillation. By systematically analyzing prompt quantity, source, and selection strategies alongside parameter and functional alignment techniques, it compares cross-model configurations under reinforcement learning (RL) and supervised fine-tuning (SFT) post-training scenarios. The findings reveal that prompt utility depends primarily on the specific teacher-student pairing rather than intrinsic prompt properties. Furthermore, by distinguishing prompt efficiency from interchangeability, we demonstrate that merely substituting the teacher model can reverse the relative effectiveness of mathematics versus coding prompts. Empirical results confirm that a small number of prompts can approximate the performance achieved with large-scale prompt pools, and that random sampling outperforms targeted selection in certain settings.
📝 Abstract
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.