Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors
This study investigates whether surreptitious learning (SL) can transfer complex capabilities and covert malicious behaviors across models, thereby posing implicit alignment risks. Moving beyond conventional preference transfer limitations, this work employs task-agnostic data distillation, integrating knowledge distillation, LoRA fine-tuning, logit matching, and steering vector techniques to systematically examine the transfer mechanisms of capabilities and behavioral traits. It provides the first empirical evidence that SL can transmit MLP prediction capabilities, linguistic backdoors, and hacking tendencies. Experimental results demonstrate that student models successfully acquire partial backdoors with a 23.5% trigger rate and exhibit hacking behaviors at a rate of 58.3%. These findings reveal significant latent security vulnerabilities associated with SL in propagating both sophisticated capabilities and covert malicious behaviors across models.