🤖 AI Summary
This study addresses a critical blind spot in mechanistic interpretability: existing circuits successfully reproduce correct model behavior yet fail to explain errors. We systematically identify this "success-biased" limitation and propose precise error reproduction as a key criterion for circuit validity. Through ablation experiments and intervention tracing on GPT-2's Indirect Object Identification task, we quantitatively demonstrate that most circuits account for only 11%–42% of model errors. By recovering omitted critical attention heads, the error reproduction rate increases to 75.1% with negligible degradation in correct prediction accuracy. This work transcends the traditional accuracy-centric paradigm, offering a novel pathway toward enhancing circuit completeness and explanatory power in mechanistic interpretability research.
📝 Abstract
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.