Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in mechanistic interpretability, where intervention-based faithfulness evaluation objectives tend to favor circuits with poor behavioral replication, resulting in an "objective-level recovery gap." This work formally defines and quantifies this concept for the first time, revealing the misranking deficiencies of metrics such as KL divergence in automated circuit discovery. Employing methods including EAP and ACDC alongside the InterpBench benchmark, the research proposes corrective strategies that leverage controlled reference editing and signal restoration techniques to mitigate context distortion. Experimental results demonstrate that these strategies successfully rectify 96% of persistent KL divergence misrankings on validation sets and held-out prompts without altering the original circuits or their behavioral scores. These findings underscore the necessity of refining evaluation objectives to effectively resolve context distortion in circuit analysis.
📝 Abstract
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mechanistic Interpretability
Circuit Discovery
Faithfulness Evaluation
Context Distortion
Objective-Level Recovery Gap