Why Subliminal Learning Needs So Much Data: A Noisy Inverse View through Steering Vector Recovery

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inefficient recovery of vector distillation signals under hard-label supervision, which forces subliminal learning to rely on massive datasets. We formulate vector distillation as a noisy linear inverse problem and leverage the Fisher information matrix to quantify the de-damping effects and noise amplification mechanisms inherent in gradient-based optimization. By employing residual stream steering vectors alongside iterative gradient descent, we systematically compare KL- and NLL-based soft versus hard supervision. Our analysis reveals that the primary role of large-scale data is to suppress label noise amplified by the Fisher inverse rather than to facilitate feature extraction. Experiments on Qwen and Gemma models demonstrate that soft supervision enables efficient feature recovery using merely hundreds of samples, thereby clarifying the true function of data scale in knowledge distillation.
πŸ“ Abstract
Subliminal learning lets a student inherit a teacher's behavioral trait from semantically unrelated data, yet published demonstrations typically require tens of thousands of carrier examples. We ask where this data requirement comes from. Our testbed is subliminal steering: the teacher trait is a known residual-stream vector $\Delta_T$, so transfer can be measured directly as parameter recovery. On identical carrier prefixes, we compare token-level (hard) NLL supervision with full-distribution (soft) KL supervision. At initialization the two objectives give nearly collinear gradients, and both align poorly with $\Delta_T$. Under iterative optimization, however, they diverge: soft supervision recovers $\Delta_T$ almost exactly from a few hundred carriers, while hard supervision stays well below it even with tens of thousands. We explain this gap by casting steering-vector distillation as a noisy linear inverse problem. Locally, the carrier task maps the trait through its Fisher matrix $F$, so gradients point toward $F\Delta_T$ rather than $\Delta_T$. Gradient descent then acts as a progressively less-damped inverse of $F$. With soft targets, this inverse restores low-curvature directions. With hard labels, it also amplifies the sampling noise in those same directions. The result is an optimal inversion depth that grows with the number of independent carriers. Experiments on Qwen2.5-7B and Gemma-2-9B confirm four predictions: the Fisher distortion of the initial gradient, recovery ordered from steep to flat directions, an optimal depth that shifts with data scale, and the finding that resampling completions from a fixed prompt pool works as well as adding new prompts. In this setting, large carrier datasets are needed less to reveal the trait than to suppress label noise amplified by Fisher inversion. Code is available at \url{https://github.com/luoyuchenmlcv/subliminal-data}.
Problem

Research questions and friction points this paper is trying to address.

Subliminal Learning
Data Efficiency
Steering Vector Recovery
Noisy Inverse Problem
Knowledge Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subliminal Learning
Steering Vector Recovery
Noisy Linear Inverse Problem
Fisher Information Matrix
Knowledge Distillation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
L
Luoyu Chen
University of Technology Sydney, Sydney, NSW, Australia
Xiaoyu Ding
Xiaoyu Ding
Research Associate, Carnegie Mellon University
Computer VisionFacial Expression AnalysisAction Unit DetectionVideo Event Detection
W
Weiqi Wang
Xi’an Jiaotong University, Xi’an, Shaanxi, China
Chenhan Zhang
Chenhan Zhang
PhD
deep Learningprivacy-preserving
Z
Zhiyi Tian
Southeast University, Nanjing, Jiangsu, China
J
Jianhuan Huang
University of Technology Sydney, Sydney, NSW, Australia
S
Shui Yu
University of Technology Sydney, Sydney, NSW, Australia