Can Data Attribution Filter Out Subliminal Learning? Not Reliably

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究评估了三种基于梯度的数据归因方法在过滤潜意识学习上的有效性,发现EK-FAC在某些情况下有效,但整体上不如基准方法。
📝 Abstract
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them against divergence tokens, a strong baseline previously shown to localize subliminal learning (albeit one that requires access to counterfactual teacher models). Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all mostly fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model-preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.
Problem

Research questions and friction points this paper is trying to address.

subliminal learning
data attribution
gradient-based attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient-based attribution
subliminal learning
data filtering
EK-FAC
divergence tokens
🔎 Similar Papers
2024-06-16International Conference on Learning RepresentationsCitations: 12