Automated Circuit Interpretation via Probe Prompting

📅 2025-11-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Neural network interpretability suffers from labor-intensive manual analysis of attribution maps—requiring ~2 hours per prompt—and lacks scalability. Method: We propose an automated subgraph extraction framework: (1) a concept probe generates high-impact features; (2) cross-prompt activation clustering and semantically aligned supernode construction identify recurrent structural motifs; and (3) three transparent, interpretability-driven rules—Semantic, Relationship, and Say-X—enable hierarchical semantic clustering. Results: Our method reveals a layered circuit structure: early Transformer layers host general-purpose mechanisms, while later layers specialize. Experiments show strong fidelity—0.83 average completeness and 0.762 activation similarity—while achieving a subgraph replacement score of 0.54 and concept-group consistency of 0.425, significantly outperforming geometric clustering baselines. The approach supports plug-and-play reproducibility, substantially improving both efficiency and trustworthiness of interpretability analysis.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Mechanistic interpretability aims to understand neural networks by identifying which learned features mediate specific behaviors. Attribution graphs reveal these feature pathways, but interpreting them requires extensive manual analysis -- a single prompt can take approximately 2 hours for an experienced circuit tracer. We present probe prompting, an automated pipeline that transforms attribution graphs into compact, interpretable subgraphs built from concept-aligned supernodes. Starting from a seed prompt and target logit, we select high-influence features, generate concept-targeted yet context-varying probes, and group features by cross-prompt activation signatures into Semantic, Relationship, and Say-X categories using transparent decision rules. Across five prompts including classic"capitals"circuits, probe-prompted subgraphs preserve high explanatory coverage while compressing complexity (Completeness 0.83, mean across circuits; Replacement 0.54). Compared to geometric clustering baselines, concept-aligned groups exhibit higher behavioral coherence: 2.3x higher peak-token consistency (0.425 vs 0.183) and 5.8x higher activation-pattern similarity (0.762 vs 0.130), despite lower geometric compactness. Entity-swap tests reveal a layerwise hierarchy: early-layer features transfer robustly (64% transfer rate, mean layer 6.3), while late-layer Say-X features specialize for output promotion (mean layer 16.4), supporting a backbone-and-specialization view of transformer computation. We release code (https://github.com/peppinob-ol/attribution-graph-probing), an interactive demo (https://huggingface.co/spaces/Peppinob/attribution-graph-probing), and minimal artifacts enabling immediate reproduction and community adoption.
Problem

Research questions and friction points this paper is trying to address.

Automating interpretation of neural network feature pathways to reduce manual analysis time
Transforming attribution graphs into compact interpretable subgraphs using concept-aligned supernodes
Grouping features by cross-prompt activation signatures to reveal behavioral coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated pipeline transforms attribution graphs into subgraphs
Groups features by cross-prompt activation signatures
Uses concept-targeted probes to identify high-influence features
🔎 Similar Papers
No similar papers found.