🤖 AI Summary
This study addresses the limitation of existing knowledge transfer methods, which typically rely on target data or teacher parameters and thus struggle to achieve efficient cross-task capability transfer. To overcome this, we propose an active task-free distillation paradigm that leverages a common ancestor model to identify nearly unrelated word pairs for constructing single-word prompts. Through unsupervised knowledge distillation, this approach uncovers "behavioral shadows" within irrelevant texts, enabling the transmission of model capabilities using only isolated words under fully decoupled conditions. Crucially, the proposed method requires neither target data nor teacher parameters. Empirical results demonstrate significant performance improvements in tasks such as code generation and logical reasoning. Furthermore, the approach exhibits strong scalability across model generations and scales, offering a novel pathway for transferring model capabilities.
📝 Abstract
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.