π€ AI Summary
This study addresses the fragmentation of visual concepts in existing sparse autoencoders (SAEs) caused by neglecting image spatial dependencies. To this end, we propose the Markov random field linear representation hypothesis, which incorporates spatial dependency priors to refine the conventional independence assumption. Building upon this, we design Spatial-SAE as an amortized maximum a posteriori (MAP) estimator to recover spatial correlations. By integrating Markov random field modeling with SAEs, our approach achieves more precise disentanglement and interpretation of visual concepts. Experimental results demonstrate that Spatial-SAE attains a 96% success rate on synthetic concept recovery tasks and significantly enhances the interpretability of DINOv2 activations.
π Abstract
Sparse Autoencoders (SAEs) decompose model activations into sparse combinations of interpretable dictionary atoms. Although SAEs are grounded in the Linear Representation Hypothesis (LRH), their objective smuggles in an additional prior: concepts across patches are treated as independent, an assumption clearly violated by natural images and by the activations they induce. We therefore specialize LRH to vision through the Markov-Field Linear Representation Hypothesis (MFLRH), which adds the missing spatial dependencies to the LRH assumptions. We thus propose Spatial-SAE as an amortized MAP estimator under the MFLRH. Spatial-SAE consistently outperforms standard SAEs in concept recovery and interpretability. Across four variants, it achieves a 96% average win rate on synthetic concept recovery and improves interpretability on DINOv2 activations, at a reconstruction cost concentrated in high spatial frequencies.