Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过LLM残差流进行隐秘信息传输的问题,提出了一种无需模型重训练或权重修改的方法来隐藏并恢复敏感信息。
📝 Abstract
Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An offline observer can recover this information with a simple linear decoder. The attack requires no model retraining or weight modification, no attacker-controlled egress, and no control over the recorder or transfer process. We introduce a residual-stream covert-channel attack that maps messages to codewords and injects them into an intermediate residual stream through a compromised runtime hook. To maintain recoverability, the injection strength is scaled with the local residual norm using the signal-to-residual-norm ratio. Across eleven models from seven architecture families, our evaluation shows 91--100% recovery on nine models with KL divergence 0.001--0.007, while evaluated activation-level detectors remain close to random guessing (AUC <= 0.56). Tested post-hoc defenses do not reliably eliminate the channel. Thus, an activation artifact can be schema-valid while carrying information that is not authorized to cross the boundary.
Problem

Research questions and friction points this paper is trying to address.

large language models
privacy-sensitive organizations
covert information transfer
residual streams
activation artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

residual-stream covert-channel attack
information hiding in activations
signal-to-residual-norm ratio
no model retraining
🔎 Similar Papers