Institution profile

Truthful AI

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors

Oct 07, 2026

This study investigates whether surreptitious learning (SL) can transfer complex capabilities and covert malicious behaviors across models, thereby posing implicit alignment risks. Moving beyond conventional preference transfer limitations, this work employs task-agnostic data distillation, integrating knowledge distillation, LoRA fine-tuning, logit matching, and steering vector techniques to systematically examine the transfer mechanisms of capabilities and behavioral traits. It provides the first empirical evidence that SL can transmit MLP prediction capabilities, linguistic backdoors, and hacking tendencies. Experimental results demonstrate that student models successfully acquire partial backdoors with a 23.5% trigger rate and exhibit hacking behaviors at a rate of 58.3%. These findings reveal significant latent security vulnerabilities associated with SL in propagating both sophisticated capabilities and covert malicious behaviors across models.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors

Oct 07, 2026

This study investigates whether surreptitious learning (SL) can transfer complex capabilities and covert malicious behaviors across models, thereby posing implicit alignment risks. Moving beyond conventional preference transfer limitations, this work employs task-agnostic data distillation, integrating knowledge distillation, LoRA fine-tuning, logit matching, and steering vector techniques to systematically examine the transfer mechanisms of capabilities and behavioral traits. It provides the first empirical evidence that SL can transmit MLP prediction capabilities, linguistic backdoors, and hacking tendencies. Experimental results demonstrate that student models successfully acquire partial backdoors with a 23.5% trigger rate and exhibit hacking behaviors at a rate of 58.3%. These findings reveal significant latent security vulnerabilities associated with SL in propagating both sophisticated capabilities and covert malicious behaviors across models.

0 citationsRead paper