Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability

📅 2026-07-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in existing interventional interpretability evaluations, which rely on point estimates and struggle to disentangle true causal effects from sampling or adaptive biases. The authors reformulate the problem as a causal estimation task and introduce, for the first time, an anytime-valid statistical certification framework that accommodates adaptive intervention sampling. By integrating Hoeffding-type confidence sequences with variance-adaptive betting strategies and bounded mixture importance weighting, the method yields both confidence intervals and dynamic confidence sequences for intervention fidelity. Empirical validation on MNIST abstractions and GPT-2 Small IOI circuit experiments demonstrates that the approach not only certifies high-fidelity interpretability claims and detects statistically insignificant differences between methods but also reduces certification costs by 10–30× compared to existing baselines.
📝 Abstract
Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices. We introduce Certified Interventional Fidelity (CIF), a statistical layer for interventional interpretability evaluations. CIF first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate CIF with Hoeffding-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by 10-30x in our experiments. On MNIST abstractions and GPT-2 Small IOI circuits, CIF certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.
Problem

Research questions and friction points this paper is trying to address.

mechanistic interpretability
causal claims
interventional evaluation
statistical certification
adaptive sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Certified Interventional Fidelity
anytime-valid confidence sequences
causal estimand
adaptive intervention sampling
mechanistic interpretability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Amir Asiaee
Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37232, USA