Backdooring Sparse Autoencoders

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the supply chain attack risks introduced by Sparse Autoencoders (SAEs) in model intervention. We propose a decoder-only backdoor scheme that freezes both the base language model and the encoder, achieving triggered malicious behavior in code generation scenarios through single-point injection. This work is the first to demonstrate that SAEs can independently host behavioral backdoors without modifying the underlying model, exposing their potential vulnerabilities as security-sensitive components. Experiments on the HumanEval benchmark show that the proposed backdoor achieves high attack success rates across multiple models and network layers, while standard SAE quality metrics remain largely unaffected.
📝 Abstract
Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.
Problem

Research questions and friction points this paper is trying to address.

Sparse Autoencoders
Backdoor Attack
Supply-chain Security
Language Models
Adversarial Machine Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Backdoor Attack
Supply-chain Security
Decoder-only Backdoor
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Enrico Ahlers
Humboldt-Universität zu Berlin
D
Daniel Passon
Humboldt-Universität zu Berlin
T
Tobias Kiecker
Humboldt-Universität zu Berlin
E
Eik Reichmann
Humboldt-Universität zu Berlin
Lars Grunske
Lars Grunske
Software Engineering, Humboldt-Universität zu Berlin, Germany
Automated Software EngineeringSafety EngineeringReliability EngineeringFormal Methods