Score
Designs, builds, and analyzes autoencoder architectures and sparse dictionary models that use expander-graph structured masks (expander dictionaries / expander SAEs) to impose and recover sparse codes. This includes constructing encoder/decoder mappings where encoder support is tied to a decoder expander mask, choosing expander masks to reduce decoder parameters to o(dn), and evaluating reconstruction fidelity of the resulting sparse representations.
This work addresses the excessive parameter count and memory overhead in conventional sparse autoencoders caused by dense decoders, which hinder efficient mechanistic interpretability. The authors propose the Expander Sparse Autoencoder, which introduces a left-$d$-regular expander graph mask into both the decoder and its tied encoder to enforce a sparse connectivity structure. This design substantially reduces learnable parameters while preserving the scale of the sparse coding problem. Theoretical analysis establishes that under expansion and column-flatness conditions, the method guarantees identifiability of $k$-sparse codes. Experiments on the Qwen2.5-3B model demonstrate that with $d=7$, the decoder’s parameters are reduced by a factor of 293 while retaining 84% of the dense model’s cross-entropy loss recovery capability, achieving a significantly improved trade-off between memory efficiency and reconstruction fidelity.
While model compression techniques (e.g., pruning, quantization) improve the inference efficiency of large language models (LLMs), their impact on interpretability—particularly via sparse autoencoders (SAEs)—remains poorly understood. Method: We systematically investigate the cross-compression-state transferability of SAEs trained on original LLMs to their pruned or quantized counterparts. We further propose structured pruning of the SAE itself—without retraining—to reduce its computational footprint while preserving explanatory capability. Results: We find that SAEs trained on uncompressed models transfer effectively to compressed variants, with only marginal degradation in explanation quality. Moreover, pruning the SAE directly achieves comparable interpretability to SAEs trained from scratch on compressed models—yet at drastically lower training cost. This work is the first to empirically validate strong SAE transferability across LLM compression states and establishes a novel, cost-efficient paradigm for maintaining interpretability in deployment-grade LLMs, thereby enabling practical, trustworthy model analysis.
Sparse autoencoders (SAEs) suffer from insufficient encoding accuracy in complex tasks. This work is the first to theoretically establish, via compressed sensing theory, an inherent and unbridgeable “amortization gap” in SAE encoders—revealing a fundamental limitation of linear–nonlinear joint encoding for sparse feature inference. Method: We propose a decoupled encoding–decoding framework that separates sparse inference from encoder learning. Specifically, we design a highly expressive learnable encoder and integrate a theoretically optimal sparse decoding algorithm with provable recovery guarantees. Results: On multiple benchmarks, our approach improves sparse code accuracy by 15–30% while maintaining near-identical computational overhead. In large language model (LLM) activation analysis, it significantly enhances both feature interpretability and localization precision. This work establishes a new paradigm bridging theoretical foundations and practical performance for SAEs.
Sparse autoencoders (SAEs) require extreme width to achieve neural interpretability, incurring prohibitive training costs. This paper proposes Switch SAE—a novel architecture that introduces sparse Mixture-of-Experts (MoE) into SAEs for the first time. It dynamically routes activations to multiple lightweight expert SAEs, enabling efficient scaling of feature capacity under fixed compute budgets. The method facilitates cross-expert feature disentanglement and sharing analysis, integrated with feature geometric modeling and interpretability evaluation. Experiments show that, under identical training budgets, Switch SAE reduces reconstruction error by up to 37% versus standard SAEs and other variants, scales feature count by over 10×, and preserves human interpretability. Its core innovation is an MoE-driven modular sparse coding paradigm, which substantially breaks the traditional trade-off between reconstruction fidelity and sparsity.
This paper identifies “feature absorption” as a critical failure mode in sparse autoencoders (SAEs) when decomposing large language model activations: monosemantic high-level features (e.g., “mathematics”) are competitively suppressed by their fine-grained subfeatures (e.g., “algebra”, “geometry”), leading to interpretability collapse. Method: We design a first-letter identification synthetic task and introduce a ground-truth–driven diagnostic framework with quantitative activation interpretability evaluation. Contribution/Results: We provide the first systematic empirical confirmation that feature absorption is pervasive, highly robust, and non-monotonically alleviated by increasing SAE scale or sparsity. Crucially, hyperparameter tuning alone cannot resolve it, necessitating foundational reformulation of SAE theory. Our work establishes the first benchmark and diagnostic paradigm for feature splitting/absorption failures in interpretable AI—offering both a standardized testbed and methodological framework for diagnosing representational pathologies in dictionary learning-based interpretability methods.
This work addresses the instability in downstream readout performance of sparse autoencoders, which can retain different linearly decodable signals despite identical reconstruction error and sparsity. To resolve this, the authors propose the Decoder-Preserving Sparse Autoencoder (DPSAE), which introduces a matrix-valued distortion metric to explicitly disentangle reconstruction quality from readout capability by embedding the optimal ridge regression predictor directly into the reconstruction loss. Furthermore, DPSAE incorporates task-prior-guided rank-relaxation optimization to modulate feature pattern selection. Evaluated on layer 8 of GPT-2 Small, DPSAE reduces held-out readout distortion by 10.6–11.4% while maintaining constant reconstruction NMSE, and its representational fidelity is confirmed through a KL non-inferiority test on natural text outputs.
This work addresses the lack of a clear theoretical characterization of the “concepts” extracted by sparse autoencoders (SAEs), particularly regarding their applicability to complex language model representations. Bypassing assumptions about data-generating models, the study extends local optimality analysis for the first time to the non-negative SAE framework. By linking non-negative jointly optimized solutions to the underlying feature distribution, it reveals how L1 regularization and non-negativity jointly shape dictionary structure. This analysis yields precise constraints between SAE features and data distribution, successfully explaining empirical phenomena such as hierarchical splitting, feature absorption, residual structures, and dense antipodal features. The results provide a theoretical foundation for SAE interpretability and offer principled guidance for future architecture design.
This work addresses the challenges of feature entanglement and degraded out-of-distribution (OOD) performance in sparse autoencoders, which arise from under-constrained training objectives and undermine interpretability. To mitigate these issues, the authors propose a mask-based regularization method that randomly replaces input tokens during training to disrupt co-occurring feature patterns. This approach effectively alleviates feature absorption, enhances the stability and robustness of latent representations, and narrows the performance gap between in-distribution and OOD settings. The method is architecture-agnostic and compatible with various sparsity levels, demonstrating consistent improvements across different sparse autoencoder configurations. Furthermore, it leads to better performance on probing tasks, indicating more disentangled and semantically meaningful representations.
This study investigates whether sparse autoencoders (SAEs) genuinely recover semantically meaningful and interpretable features in neural networks, rather than merely achieving high reconstruction accuracy. The authors introduce a synthetic data setting with known ground-truth features to quantitatively measure SAEs’ feature recovery rates for the first time, complemented by evaluations on real model activations. To rigorously assess performance, they propose three constrained random baselines based on random directions or activation patterns. Results reveal that SAEs recover only about 9% of the true features and show no statistically significant advantage over random baselines across key interpretability metrics—including interpretability scores (0.87 vs. 0.90), sparse probing (0.69 vs. 0.72), and causal editing (0.73 vs. 0.72)—challenging the presumed efficacy of SAEs in current interpretability applications.
This work addresses the instability and high reconstruction error commonly observed in sparse autoencoders (SAEs), which stem from non-identifiability leading to inconsistent dictionaries and encodings during training. To overcome this limitation, the authors propose the identifiable Sparse Autoencoder (iSAE), built upon the TopK SAE architecture. By introducing structural modifications and a training strategy that enforces an approximate restricted isometry condition, iSAE achieves— for the first time in practice—approximately identifiable sparse codes. This advancement significantly enhances model stability and reduces reconstruction error, while also establishing a theoretical bridge between modern SAEs and classical dictionary learning frameworks.