🤖 AI Summary
Neural network interpretability suffers from labor-intensive manual analysis of attribution maps—requiring ~2 hours per prompt—and lacks scalability. Method: We propose an automated subgraph extraction framework: (1) a concept probe generates high-impact features; (2) cross-prompt activation clustering and semantically aligned supernode construction identify recurrent structural motifs; and (3) three transparent, interpretability-driven rules—Semantic, Relationship, and Say-X—enable hierarchical semantic clustering. Results: Our method reveals a layered circuit structure: early Transformer layers host general-purpose mechanisms, while later layers specialize. Experiments show strong fidelity—0.83 average completeness and 0.762 activation similarity—while achieving a subgraph replacement score of 0.54 and concept-group consistency of 0.425, significantly outperforming geometric clustering baselines. The approach supports plug-and-play reproducibility, substantially improving both efficiency and trustworthiness of interpretability analysis.
📝 Abstract
Mechanistic interpretability aims to understand neural networks by identifying which learned features mediate specific behaviors. Attribution graphs reveal these feature pathways, but interpreting them requires extensive manual analysis -- a single prompt can take approximately 2 hours for an experienced circuit tracer. We present probe prompting, an automated pipeline that transforms attribution graphs into compact, interpretable subgraphs built from concept-aligned supernodes. Starting from a seed prompt and target logit, we select high-influence features, generate concept-targeted yet context-varying probes, and group features by cross-prompt activation signatures into Semantic, Relationship, and Say-X categories using transparent decision rules. Across five prompts including classic"capitals"circuits, probe-prompted subgraphs preserve high explanatory coverage while compressing complexity (Completeness 0.83, mean across circuits; Replacement 0.54). Compared to geometric clustering baselines, concept-aligned groups exhibit higher behavioral coherence: 2.3x higher peak-token consistency (0.425 vs 0.183) and 5.8x higher activation-pattern similarity (0.762 vs 0.130), despite lower geometric compactness. Entity-swap tests reveal a layerwise hierarchy: early-layer features transfer robustly (64% transfer rate, mean layer 6.3), while late-layer Say-X features specialize for output promotion (mean layer 16.4), supporting a backbone-and-specialization view of transformer computation. We release code (https://github.com/peppinob-ol/attribution-graph-probing), an interactive demo (https://huggingface.co/spaces/Peppinob/attribution-graph-probing), and minimal artifacts enabling immediate reproduction and community adoption.