🤖 AI Summary
This study addresses the imprecise behavioral localization in mechanistic interpretability of large language models caused by neglecting component interaction effects. To this end, it proposes the WISE causal estimator and JuntaLearner, a gradient-based circuit discovery method. The core innovation lies in pioneering set-effect estimation integrated with a witness mechanism, which effectively circumvents combinatorial enumeration explosion, alongside the introduction of the CRS metric for comprehensively evaluating the critical roles of small circuits. By synergizing causal inference, witness variables, and gradient optimization techniques, the proposed approach achieves scalable, high-precision circuit discovery. Experimental results demonstrate that JuntaLearner significantly outperforms baselines in average CRS across diverse tasks and varying model scales.
📝 Abstract
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the witness-integrated set effect (WISE), a family of causal estimands that build on the witness mechanism while taking expectations over sets of causes and witnesses to remain computationally feasible. Building on this approach, we introduce JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the circuit recognition score (CRS) to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, JuntaLearner achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, JuntaLearner scales to large models while accounting for set-level interactions and avoiding first-order approximations.