Score
Design and train autoencoder architectures and training procedures that produce sparse, localized latent codes and feature dictionaries shared across multiple inputs or views, using joint factorization or explicit alignment constraints to recover interpretable, monosemantic components. Build probes and intervention-capable analysis pipelines to extract, inspect, and manipulate those sparse latent directions for controlled, interpretable representations.
Addressing the fundamental opacity of large language models (LLMs), this paper introduces the first end-to-end sparse autoencoder (SAE) framework explicitly designed for LLM interpretability. Methodologically, it integrates feature disentanglement modeling, L1/L0 sparsity regularization, intermediate-layer feature distillation, and multi-dimensional interpretability evaluation—including feature consistency and semantic readability—to systematically decode latent neural activations. Key contributions include: (1) establishing the first comprehensive SAE technical stack spanning theoretical foundations, architectural variants, and behavioral intervention mechanisms; (2) shifting LLM interpretability paradigms from passive “black-box analysis” toward active “editable neuron” control; and (3) empirically demonstrating efficacy in concept localization, bias attribution, and controllable generation. The framework provides a foundational methodology for next-generation LLMs that are both interpretable and editable.
Classical sparse autoencoders (SAEs) and variational autoencoders (VAEs) suffer from structural limitations in modeling multi-manifold data—specifically, theoretical constraints under manifold union assumptions and biased intrinsic dimension estimation. Method: We propose a novel hybrid autoencoder integrating deterministic sparse regularization with stochastic encoding, augmented by manifold-aware optimization. Contribution/Results: We formally characterize, for the first time, the theoretical limitations of SAEs/VAEs under multi-manifold union structures. We prove that our model’s global optimum exactly recovers the underlying multi-manifold geometry and yields unbiased intrinsic dimension estimates. Guided by theory, we design an end-to-end loss function. Empirically, our method outperforms same-capacity SAEs, VAEs, and several diffusion models on synthetic benchmarks, real-world images, and large-language-model activation data—achieving lower reconstruction error, sparser latent representations, and more accurate manifold dimension estimation.
Current visual model interpretability research faces two key bottlenecks: (1) features lack causal validation capability, and (2) controllable editing remains infeasible without fine-tuning. This paper introduces the first unified framework based on sparse autoencoders (SAEs) that jointly enables discovery of human-interpretable visual features, disentangled representation learning, and causal intervention validation. Methodologically, we systematically integrate SAEs into internal visual model representations for both interpretability modeling and causal testing—performing targeted, activation-level interventions to conduct controlled hypothesis-testing experiments. Evaluated across multiple state-of-the-art vision models, our framework consistently identifies and manipulates semantically coherent high-level visual features (e.g., texture, parts, object attributes), with interventions demonstrating reproducibility and attribution fidelity. Code, pretrained models, and an interactive demo are publicly released.
This study addresses the challenges of latent variable identifiability and poor interpretability in high-dimensional single-cell multi-omics data. Methodologically, we systematically investigate the capacity of sparse autoencoders (SAEs) to disentangle latent variables and uncover underlying biological mechanisms. We first validate SAEs’ ability to recover ground-truth generative latents using controllable synthetic data; then perform end-to-end modeling and interpretability analysis on real single-cell multi-omics datasets. Our key contributions are: (1) the first empirical demonstration that SAEs, in a fully unsupervised setting, automatically identify biologically meaningful processes—such as CO₂ transport and ion homeostasis; (2) strong biological consistency in discovering erythroid differentiation and immune regulatory pathways; and (3) a theoretical advance overcoming the identifiability limitation of overparameterized models. Collectively, this work establishes a novel paradigm for interpretable representation learning in multi-omics.
This work addresses the issue that latent features in sparse autoencoders (SAEs) often conflate multiple semantic subspaces, compromising atomicity and interpretability. To mitigate this, the authors propose a joint training approach that introduces a small meta-SAE to sparsely reconstruct the columns of the main SAE’s decoder. They further impose a decomposability penalty on directions that are easily reconstructible by the meta dictionary, directly optimizing feature atomicity during training. This method is the first to explicitly treat decomposability as a training objective, effectively reducing cross-semantic aliasing. Experiments show a 7.5% reduction in average |φ| and a 7.6% improvement in automatic interpretability scores on GPT-2 Large, along with an 8.6% gain in Fuzz scores on Gemma 2 9B. Qualitative analysis confirms that polysemantic features are successfully disentangled into semantically coherent atomic subfeatures.
This paper identifies “feature absorption” as a critical failure mode in sparse autoencoders (SAEs) when decomposing large language model activations: monosemantic high-level features (e.g., “mathematics”) are competitively suppressed by their fine-grained subfeatures (e.g., “algebra”, “geometry”), leading to interpretability collapse. Method: We design a first-letter identification synthetic task and introduce a ground-truth–driven diagnostic framework with quantitative activation interpretability evaluation. Contribution/Results: We provide the first systematic empirical confirmation that feature absorption is pervasive, highly robust, and non-monotonically alleviated by increasing SAE scale or sparsity. Crucially, hyperparameter tuning alone cannot resolve it, necessitating foundational reformulation of SAE theory. Our work establishes the first benchmark and diagnostic paradigm for feature splitting/absorption failures in interpretable AI—offering both a standardized testbed and methodological framework for diagnosing representational pathologies in dictionary learning-based interpretability methods.
This work identifies and formally names a previously uncharacterized phenomenon in vision-language models termed “cross-modal feature heterogeneity,” wherein semantically equivalent concepts activate sparse features along inconsistent directions across modalities, leading to modality fragmentation that undermines interpretability and controllability. The study demonstrates that mere alignment of activations is insufficient to resolve this feature mismatch. To address this, the authors propose a two-stage strategy: first preserving each modality’s intrinsic feature geometry using modality-specific sparse autoencoders, followed by post-hoc alignment of corresponding cross-modal features. This approach significantly improves reconstruction fidelity and achieves superior performance in cross-modal retrieval and concept-guided generation tasks.
This work addresses the limitations of conventional sparse autoencoders, which model latent features as one-dimensional vectors and thereby fail to capture the intrinsic multidimensional structure of language model representations, often resulting in feature splitting and loss of geometric information. To overcome this, the authors propose the Subspace-Aware Sparse Autoencoder (SASA), which generalizes the decoder from a single vector to a learnable low-dimensional subspace. SASA enforces block sparsity via Top-s group gating and employs nuclear norm regularization to adaptively control the effective rank of each group. Theoretically, when the block size is at least the intrinsic feature dimension, a single group can globally reconstruct an entire feature slice with optimal fidelity, substantially reducing sample complexity. Experiments on GPT-2 and Mistral-7B demonstrate that SASA effectively mitigates feature splitting and absorption, enhances representational monosemy and interpretability, and matches or exceeds the performance of standard methods using roughly half the training tokens.
This work proposes LUCID, the first unified vision-language sparse autoencoder that addresses the limitations of modality-specific training in existing approaches, which often yield uninterpretable features and hinder cross-modal alignment. LUCID employs a shared-private representation architecture to learn a common latent dictionary across modalities while preserving modality-specific characteristics. Unsupervised feature alignment is achieved through optimal transport, enabling pixel-level localization and cross-modal neuron correspondence. This design substantially mitigates concept clustering issues and enhances interpretability. Furthermore, LUCID introduces an automated dictionary interpretation pipeline based on term clustering. The learned shared features encompass objects, actions, attributes, and abstract concepts, significantly improving both the interpretability and robustness of multimodal representations.
This work addresses the feature splitting and feature absorption phenomena commonly observed in sparse autoencoders when trained with large-scale dictionaries, which undermine the atomicity and interpretability of latent representations. To mitigate these issues, the authors propose Cross-sample Consistency Regularization (C²R), a novel regularization technique that encourages semantically similar samples within a batch to activate consistent latent units while suppressing the co-activation of latent variables with similar directions. This approach systematically enhances the discreteness and semantic clarity of learned features without compromising reconstruction fidelity. Experimental results demonstrate that C²R significantly improves the atomicity and interpretability of latent representations, offering a promising direction for advancing sparse representation learning.
This work addresses the limited semantic interpretability of embeddings in conventional latent factor recommender models, such as matrix factorization, which undermines system transparency and controllability. The authors propose the first application of Matryoshka Sparse Autoencoders (MSAE) to collaborative filtering for learning hierarchical, sparse, and disentangled latent representations. Evaluated on the Amazon Fashion dataset, the approach demonstrates strong semantic coherence of extracted features through alignment with item metadata and automatic annotation via large language models. Furthermore, neuron-level interventions—such as manipulating gender-related latent variables—validate the model’s interpretability and controllability. The method effectively mitigates feature splitting and composition issues commonly observed in traditional sparse autoencoders when scaled to larger embedding dimensions.