🤖 AI Summary
This work addresses the limited interpretability of internal representations in large language models, a challenge exacerbated by existing post-hoc methods like sparse autoencoders (SAEs), which incur substantial computational overhead. The authors propose ParityTransformer, which integrates a Deep Parity Bottleneck (DPB) into a GPT-2–scale model to natively induce layer-wise, wide-dimensional, and sparse intermediate representations within the Transformer architecture for the first time. By combining a parameter-free algebraic dictionary—ensuring deterministic incoherence—with a multi-level mixture-of-experts structure to efficiently enforce sparsity, the approach significantly reduces both computational and memory costs. Experiments demonstrate that ParityTransformer matches SAE performance on sparse probing tasks while outperforming it in feature absorption, intervention steering, and fine-grained causal mediation, confirming that the model can directly leverage interpretable features for reasoning.
📝 Abstract
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction has remained impractical, as a per-layer over-complete bottleneck is prohibitively expensive in both memory and compute. To overcome this issue, we introduce the ParityTransformer, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design. At each layer, a Deep Parity Bottleneck (DPB) replaces a learned over-complete basis with a parameter-free algebraic dictionary, providing a deterministic incoherence guarantee and eliminating the memory requirements that have prevented per-layer interpretable bottlenecks at scale. A DPB is a hierarchically structured sparse bottleneck which efficiently enforces sparsity using a multi-level mixture-of-experts approach: a hardware-aware implementation that closes the cost gap between activation sparse and dense training to a manageable interpretability tax. Empirically, ParityTransformers perform at least as well as post-hoc SAEs on sparse probing tasks, while out-performing on measures of feature absorption, steering effectiveness, and fine-grained causal interventions. Because subsequent computation acts only on features that survive the sparse bottleneck, the ParityTransformer's features are native to the model's forwards pass by construction, addressing the question of whether SAEs probe features the model actually uses during computation. We see this as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.