π€ AI Summary
This study addresses the limitations of conventional knowledge distillation, which relies on parameter updates that bind knowledge to specific models and hinder reusability, rendering it impractical for API-restricted or cost-prohibitive scenarios. To overcome this, the authors propose a universal, parameter-free text teaching framework that translates teacher model knowledge into interpretable, reusable natural language artifacts termed βPrimers.β Through a multi-agent collaboration mechanism encompassing student trial-and-error, prompt transformation, teacher demonstration, and synthetic integration, these architecture-agnostic textual artifacts are iteratively optimized. The proposed approach enables zero-parameter knowledge injection across diverse models. Empirical evaluations on benchmarks such as Omni-MATH demonstrate substantial performance gains, achieving 51.7% accuracy in mathematical reasoning and outperforming mainstream prompt engineering techniques as well as parameterized distillation methods.
π Abstract
Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer. Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student's accuracy from 9.4% to 48.6% and Fast1 accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher-Student pair can generalize to other Students that do not participate in the synthesis.