🤖 AI Summary
This study addresses the challenges of unverified teacher responses and shared error reinforcement in unlabeled knowledge distillation by proposing a novel framework that requires neither weight updates nor ground-truth labels. Methodologically, the approach synthesizes prompts, refines paired responses, and filters candidates through an answer-consistency-based adaptive search. Theoretically, it decouples the generation and selection processes while establishing formal conditions for accuracy guarantees. Experimental results demonstrate that the proposed method significantly outperforms existing unlabeled distillation techniques on reasoning tasks, achieving performance comparable to supervised optimization while requiring only the deployment of a frozen student model.
📝 Abstract
Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.