Inference and learning in sparse autoencoders as natural gradient flow

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of reliably recovering interpretable features using sparse autoencoders (SAEs) under feature superposition and low-frequency activation. To this end, we propose BeFOND, a model that unifies inference and dictionary learning into a natural gradient flow, enabling encoder-free sparse coding. Specifically, BeFOND eliminates feature interference through recurrent inference and compensates for the slow learning of rare features via Fisher information matrix preconditioning. Experimental results demonstrate that our approach significantly enhances both dictionary recovery and rare feature detection. Notably, BeFOND outperforms pretrained SAEs with substantially less data, and its performance continues to improve consistently as dictionary width scales up.
πŸ“ Abstract
Sparse autoencoders are widely used to uncover interpretable features in neural networks, yet reliable recovery remains difficult when features overlap or activate infrequently. These challenges involve both inferring which features explain an input and learning the dictionary that represents them. Here, we unify inference and dictionary learning as natural-gradient flows on a shared variational free energy. We instantiate this framework as BeFOND, an encoder-free sparse coding model with closed-form inference and learning dynamics. We show how recurrent explaining away reduces interference between overlapping features, while Fisher preconditioning can compensate for the slow learning of rare features. On synthetic data, BeFOND improves dictionary recovery and rare-feature detection, with a growing advantage over amortized baselines as superposition increases. On language-model activations, it improves single-feature concept detection and selective intervention, outperforming pretrained reference SAEs with substantially less training data. Its feature quality continues to improve with dictionary width, whereas the evaluated baselines largely plateau. Together, these results show how improving inference and learning within a unified probabilistic framework can make better use of data and dictionary capacity to interpret and intervene on neural representations.
Problem

Research questions and friction points this paper is trying to address.

sparse autoencoders
feature recovery
dictionary learning
overlapping features
rare features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Natural Gradient Flow
Encoder-free Sparse Coding
Fisher Preconditioning
Recurrent Explaining Away
πŸ”Ž Similar Papers