A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models

📅 2025-03-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Addressing the fundamental opacity of large language models (LLMs), this paper introduces the first end-to-end sparse autoencoder (SAE) framework explicitly designed for LLM interpretability. Methodologically, it integrates feature disentanglement modeling, L1/L0 sparsity regularization, intermediate-layer feature distillation, and multi-dimensional interpretability evaluation—including feature consistency and semantic readability—to systematically decode latent neural activations. Key contributions include: (1) establishing the first comprehensive SAE technical stack spanning theoretical foundations, architectural variants, and behavioral intervention mechanisms; (2) shifting LLM interpretability paradigms from passive “black-box analysis” toward active “editable neuron” control; and (3) empirically demonstrating efficacy in concept localization, bias attribution, and controllable generation. The framework provides a foundational methodology for next-generation LLMs that are both interpretable and editable.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large Language Models (LLMs) have revolutionized natural language processing, yet their internal mechanisms remain largely opaque. Recently, mechanistic interpretability has attracted significant attention from the research community as a means to understand the inner workings of LLMs. Among various mechanistic interpretability approaches, Sparse Autoencoders (SAEs) have emerged as a particularly promising method due to their ability to disentangle the complex, superimposed features within LLMs into more interpretable components. This paper presents a comprehensive examination of SAEs as a promising approach to interpreting and understanding LLMs. We provide a systematic overview of SAE principles, architectures, and applications specifically tailored for LLM analysis, covering theoretical foundations, implementation strategies, and recent developments in sparsity mechanisms. We also explore how SAEs can be leveraged to explain the internal workings of LLMs, steer model behaviors in desired directions, and develop more transparent training methodologies for future models. Despite the challenges that remain around SAE implementation and scaling, they continue to provide valuable tools for understanding the internal mechanisms of large language models.
Problem

Research questions and friction points this paper is trying to address.

Understanding opaque internal mechanisms of Large Language Models (LLMs).
Exploring Sparse Autoencoders (SAEs) for interpretable feature disentanglement in LLMs.
Developing transparent training methodologies and steering LLM behaviors using SAEs.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders disentangle LLM features
SAEs explain and steer LLM behaviors
SAEs enhance LLM training transparency