🤖 AI Summary
This study addresses a critical gap in current mechanistic interpretability research, where causal claims are often made without explicit articulation of identification assumptions, conflating validation metrics with genuine causal identification. The authors conduct the first systematic audit of 40 representative papers, employing dual-coder annotation and qualitative content analysis to evaluate the alignment between stated causal claims and underlying identification assumptions. Findings reveal that the vast majority of papers omit dedicated statements of identification assumptions and routinely substitute validation for identification. To rectify this, the paper proposes a standardized disclosure framework encompassing explicit assumption statements, clear naming of identification strategies, and sensitivity analyses, thereby advocating for more rigorous standards of causal inference in the field.
📝 Abstract
Mechanistic interpretability papers increasingly use causal vocabulary: circuits, mediators, causal abstraction, monosemanticity. Such claims require explicit identification assumptions. A purposive audit of 10 papers across four methodological strands finds no dedicated identification-assumptions section and a recurring pattern: validation metrics such as faithfulness, completeness, monosemanticity, alignment, or ablation effects are reported as causal support without stating the assumptions that make them identifying. A two-human-coder audit on $n=30$ reproduces the direction of the main finding: dedicated identification sections are absent, and validation-metric substitution is common, though exact Dim B/D counts are coding-rule sensitive. The paper proposes a disclosure norm: state whether the claim is causal, name the identification strategy, enumerate assumptions, stress at least one, and explain how conclusions shift if assumptions fail. Validation is not identification.