Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过学习动态和跨任务核函数方法,解释了潜意识学习中学生模型如何在没有直接标签的情况下从教师模型的辅助输出中获得任务能力。
📝 Abstract
Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first learning-dynamics-based theoretical explanation of how ghost-output supervision gives rise to subliminal learning.
Problem

Research questions and friction points this paper is trying to address.

Subliminal Learning
Optimization Mechanism
Ghost Outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

chained cross-task kernel
shared backbone representations
Positive Semi-Definite (PSD) structure
rank bottleneck
high-entropy inputs
🔎 Similar Papers
No similar papers found.