Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of prompt data on teacher-student knowledge transfer in online policy distillation. By systematically analyzing prompt quantity, source, and selection strategies alongside parameter and functional alignment techniques, it compares cross-model configurations under reinforcement learning (RL) and supervised fine-tuning (SFT) post-training scenarios. The findings reveal that prompt utility depends primarily on the specific teacher-student pairing rather than intrinsic prompt properties. Furthermore, by distinguishing prompt efficiency from interchangeability, we demonstrate that merely substituting the teacher model can reverse the relative effectiveness of mathematics versus coding prompts. Empirical results confirm that a small number of prompts can approximate the performance achieved with large-scale prompt pools, and that random sampling outperforms targeted selection in certain settings.
📝 Abstract
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
knowledge transfer
prompt selection
data efficiency
teacher-student alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Prompt efficiency
Knowledge transfer
Functional alignment
Data selection
🔎 Similar Papers
No similar papers found.
Jiaxuan Wang
Jiaxuan Wang
GE HealthCare
Model interpretabilityMachine learning for healthcareOut of distribution generalization
Jiafei Lyu
Jiafei Lyu
PhD of Control Science and Engineering, Tsinghua University
deep reinforcement learning
Y
Yuchen Cai
LLM Department, Tencent
Siye Wu
Siye Wu
Fudan University
P
Pengyuan Wang
LLM Department, Tencent
J
Jiashun Liu
LLM Department, Tencent
X
Xiang Cheng
Renmin University of China
K
Kai Yang
LLM Department, Tencent
Y
Yangkun Chen
LLM Department, Tencent
S
Saiyong Yang
LLM Department, Tencent
Lan-Zhe Guo
Lan-Zhe Guo
LAMDA Group, Nanjing University
Machine Learning