MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents

📅 2026-04-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the issue that latent features in sparse autoencoders (SAEs) often conflate multiple semantic subspaces, compromising atomicity and interpretability. To mitigate this, the authors propose a joint training approach that introduces a small meta-SAE to sparsely reconstruct the columns of the main SAE’s decoder. They further impose a decomposability penalty on directions that are easily reconstructible by the meta dictionary, directly optimizing feature atomicity during training. This method is the first to explicitly treat decomposability as a training objective, effectively reducing cross-semantic aliasing. Experiments show a 7.5% reduction in average |φ| and a 7.6% improvement in automatic interpretability scores on GPT-2 Large, along with an 8.6% gain in Fuzz scores on Gemma 2 9B. Qualitative analysis confirms that polysemantic features are successfully disentangled into semantically coherent atomic subfeatures.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Safety and RobustnessSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Sparse autoencoders (SAEs) are increasingly used for safety-relevant applications including alignment detection and model steering. These use cases require SAE latents to be as atomic as possible. Each latent should represent a single coherent concept drawn from a single underlying representational subspace. In practice, SAE latents blend representational subspaces together. A single feature can activate across semantically distinct contexts that share no true common representation, muddying an already complex picture of model computation. We introduce a joint training objective that directly penalizes this subspace blending. A small meta SAE is trained alongside the primary SAE to sparsely reconstruct the primary SAE's decoder columns; the primary SAE is penalized whenever its decoder directions are easy to reconstruct from the meta dictionary. This occurs whenever latent directions lie in a subspace spanned by other primary directions. This creates gradient pressure toward more mutually independent decoder directions that resist sparse meta-compression. On GPT-2 large (layer 20), the selected configuration reduces mean $|\varphi|$ by 7.5% relative to an identical solo SAE trained on the same data. Automated interpretability (fuzzing) scores improve by 7.6%, providing external validation of the atomicity gain independent of the training and co-occurrence metrics. Reconstruction overhead is modest. Results on Gemma 2 9B are directional. On not-fully-converged SAEs, the same parameterization yields the best results, a $+8.6\%$ $Δ$Fuzz. Though directional, this is an encouraging sign that the method transfers to a larger model. Qualitative analysis confirms that features firing on polysemantic tokens are split into semantically distinct sub-features, each specializing in a distinct representational subspace.
Problem

Research questions and friction points this paper is trying to address.

sparse autoencoders
atomicity
subspace blending
polysemantic features
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Atomicity
Subspace Decomposability
Joint Training
Interpretability