Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

📅 2026-07-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the instability in downstream readout performance of sparse autoencoders, which can retain different linearly decodable signals despite identical reconstruction error and sparsity. To resolve this, the authors propose the Decoder-Preserving Sparse Autoencoder (DPSAE), which introduces a matrix-valued distortion metric to explicitly disentangle reconstruction quality from readout capability by embedding the optimal ridge regression predictor directly into the reconstruction loss. Furthermore, DPSAE incorporates task-prior-guided rank-relaxation optimization to modulate feature pattern selection. Evaluated on layer 8 of GPT-2 Small, DPSAE reduces held-out readout distortion by 10.6–11.4% while maintaining constant reconstruction NMSE, and its representational fidelity is confirmed through a KL non-inferiority test on natural text outputs.
📝 Abstract
Sparse autoencoders (SAEs) compress model activations into sparse codes, but equal reconstruction error and sparsity can preserve different linearly decodable signals. We formalize this ambiguity as a matrix-valued distortion between optimal ridge-prediction operators and train decoder-preserving SAEs by combining this distortion with reconstruction loss. In a rank relaxation, an isotropic task prior saturates per-mode omission costs without changing PCA's ordering, whereas a structured prior can change which modes are retained. A controlled sparse experiment shows that a declared prior protects held-out combinations from its task subspace. On GPT-2 small block 8, DPSAE reduces held-out decoder distortion by 10.6--11.4% across three paired runs while matching reconstruction NMSE. The same checkpoints pass an average natural-text output-KL noninferiority test, but one matched Pythia pair shows no improvement in probes restricted to a few sparse features. These results show that reconstruction quality does not determine which refitted linear readouts survive sparse compression, and that readout preservation is distinct from learning cleaner benchmark concepts or preserving every frozen-model behavior.
Problem

Research questions and friction points this paper is trying to address.

sparse autoencoders
linear readouts
decoder preservation
reconstruction error
sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

decoder-preserving
sparse autoencoders
linear readouts
matrix-valued distortion
task prior
🔎 Similar Papers
No similar papers found.