🤖 AI Summary
Existing sparse autoencoders are constrained by single-parent tree topology assumptions, limiting their ability to reliably recover multi-parent mixed feature relationships. This work proposes an adaptive graph sparse autoencoder that introduces complete parent sets as atomic hypotheses within a competitive verification mechanism. By integrating differentiable structural losses with persistent reconstruction gap initialization to guide training, the method establishes a self-consistent iterative optimization loop between dictionary learning and graph structure discovery, thereby overcoming conventional topological limitations. Experiments demonstrate that this approach achieves precise recovery of mixed topologies in controlled settings and significantly enhances relational reliability, semantic validity, and causal intervention capabilities when applied to real large language model activations.
📝 Abstract
Sparse autoencoders (SAEs) expose interpretable features in large language model activations, yet existing structured SAEs impose single-parent trees or forests, while post-hoc graphs permit multiple parents but neither guide feature learning nor ensure reliable relation recovery. We introduce the Adaptive Graph Sparse Autoencoder (AG-SAE), a structure-guided training paradigm that treats each feature's complete parent set as an atomic structural hypothesis and lets evidence select zero, one, or multiple parents. By competing complete parent sets against null, subset, and alternative explanations, AG-SAE identifies jointly necessary multi-parent relations while rejecting redundant or spurious alternatives and verifying that each child contributes beyond its parents. The induced topology over SAE features then defines a differentiable structural loss that guides SAE training, while topology-guided refinement mitigates feature absorption and uses persistent reconstruction gaps exposed by the learned structure to initialize new features. The entire graph is then induced again from the revised dictionary by reassessing every feature's complete parent set, closing the dictionary-graph self-consistency cycle. Experiments demonstrate exact mixed-topology recovery in a controlled toy model, greater relational reliability and semantic validity than structured and post-hoc baselines on real LLM activations, and stronger feature-level causal interventions than conventional SAE features. AG-SAE thereby turns recovered mixed-topology feature structure into an unsupervised training signal that improves the dictionary, enables reliable feature organization beyond the topological limitations of trees, and exhibits stronger causal control beyond reconstruction.