A Testable Theory of Atomic Features

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability and scale-dependent fragmentation of sparse autoencoder (SAE) features in language model representations by proposing an "atomic feature" theoretical framework. By training and analyzing SAEs of varying scales on their capacity to decompose large embedding models, this work uncovers intrinsic patterns governing feature evolution. The results demonstrate that as model scale increases, sparse dictionaries can stably recover high-frequency atomic features. Furthermore, SAE features exhibit significant stability across different datasets and scales, along with hierarchical inclusion relationships. These findings challenge conventional assumptions regarding feature fragmentation. Through experimental validation of feature recovery principles, this research provides critical empirical evidence for establishing a representational science grounded in atomic features.
📝 Abstract
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This"recovery principle"yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and"split"as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.
Problem

Research questions and friction points this paper is trying to address.

atomic features
sparse autoencoders
language model representations
feature stability
scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Atomic Features
Sparse Autoencoders
Recovery Principle
Feature Stability
Representation Theory