Post-Training Leaves Behavioral Shadows on Unrelated Decisions

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing knowledge transfer methods, which typically rely on target data or teacher parameters and thus struggle to achieve efficient cross-task capability transfer. To overcome this, we propose an active task-free distillation paradigm that leverages a common ancestor model to identify nearly unrelated word pairs for constructing single-word prompts. Through unsupervised knowledge distillation, this approach uncovers "behavioral shadows" within irrelevant texts, enabling the transmission of model capabilities using only isolated words under fully decoupled conditions. Crucially, the proposed method requires neither target data nor teacher parameters. Empirical results demonstrate significant performance improvements in tasks such as code generation and logical reasoning. Furthermore, the approach exhibits strong scalability across model generations and scales, offering a novel pathway for transferring model capabilities.
📝 Abstract
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
Problem

Research questions and friction points this paper is trying to address.

post-training
capability transfer
language models
knowledge distillation
behavioral shadow
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Taskless Distillation
Behavioral Shadows
Capability Transfer
Subliminal Learning
Post-Training
🔎 Similar Papers
No similar papers found.
Z
Ziyang Zhang
Peking University
Y
Yubin Jing
Georgia Institute of Technology
Y
Yuanhao Zeng
ShanghaiTech University
Y
Yuyao Li
Tsinghua University
Haofan Wang
Haofan Wang
Lovart AI, InstantX, Carnegie Mellon University
AI DesignImage GenerationGenerative AI
Y
Yichen Gong
Lovart AI