From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of implicit functional connectivity and high evidence-gathering costs in interpreting sparse autoencoder features by proposing DAFI, an intelligent agent framework. This method introduces an end-to-end functional interpretation paradigm that achieves comprehensive analysis from input semantics to output effects through active evidence collection and feedback-driven optimization. Furthermore, it employs short-context token probing to circumvent full-corpus scanning and distills successful skills to enhance efficiency. Experimental results demonstrate that DAFI increases the joint pass rate to 92%, significantly outperforming baseline methods. By substantially reducing token consumption while effectively identifying non-equivalent semantic features, this work provides a highly efficient new paradigm for mechanistic interpretability.
📝 Abstract
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
Problem

Research questions and friction points this paper is trying to address.

Sparse Autoencoders
Mechanistic Interpretability
Feature Interpretation
Functional Interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Mechanistic Interpretability
Functional Interpretation
Agentic Feature Interpretation
Skill Distillation
🔎 Similar Papers
2024-04-22International Conference on Machine LearningCitations: 10