🤖 AI Summary
This study addresses the challenges of implicit functional connectivity and high evidence-gathering costs in interpreting sparse autoencoder features by proposing DAFI, an intelligent agent framework. This method introduces an end-to-end functional interpretation paradigm that achieves comprehensive analysis from input semantics to output effects through active evidence collection and feedback-driven optimization. Furthermore, it employs short-context token probing to circumvent full-corpus scanning and distills successful skills to enhance efficiency. Experimental results demonstrate that DAFI increases the joint pass rate to 92%, significantly outperforming baseline methods. By substantially reducing token consumption while effectively identifying non-equivalent semantic features, this work provides a highly efficient new paradigm for mechanistic interpretability.
📝 Abstract
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.