🤖 AI Summary
This study addresses the limitation of existing methods that treat activation and parameter spaces in isolation, lacking interpretable causal editing capabilities. To this end, we propose Activation-Supported Parameter Decomposition (ASPD), a framework that jointly decomposes both spaces. By introducing an internal reconstruction objective to anchor weight components to activation features—effectively constraining weights via activations—ASPD resolves the non-uniqueness inherent in parameter decomposition while enabling causal editing. Experiments on Qwen-3-8B successfully recover the Indirect Object Identification (IOI) circuit mechanism and trace semantic transformations, validating the scalability and interpretability of our approach in large language models.
📝 Abstract
Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model's internal activations, while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed. Together, these properties enable scalable, interpretable, and causally editable parameter decomposition in pretrained large language models, demonstrated on Qwen-3-8B. The learned read--write components can also be composed into parameter-level mechanism circuits. We use ASPD to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights.