Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imprecise behavioral localization in mechanistic interpretability of large language models caused by neglecting component interaction effects. To this end, it proposes the WISE causal estimator and JuntaLearner, a gradient-based circuit discovery method. The core innovation lies in pioneering set-effect estimation integrated with a witness mechanism, which effectively circumvents combinatorial enumeration explosion, alongside the introduction of the CRS metric for comprehensively evaluating the critical roles of small circuits. By synergizing causal inference, witness variables, and gradient optimization techniques, the proposed approach achieves scalable, high-precision circuit discovery. Experimental results demonstrate that JuntaLearner significantly outperforms baselines in average CRS across diverse tasks and varying model scales.
📝 Abstract
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the witness-integrated set effect (WISE), a family of causal estimands that build on the witness mechanism while taking expectations over sets of causes and witnesses to remain computationally feasible. Building on this approach, we introduce JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the circuit recognition score (CRS) to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, JuntaLearner achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, JuntaLearner scales to large models while accounting for set-level interactions and avoiding first-order approximations.
Problem

Research questions and friction points this paper is trying to address.

mechanistic interpretability
circuit discovery
actual causality
language models
context-dependent effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

mechanistic interpretability
actual causality
circuit discovery
witness-integrated set effect (WISE)
JuntaLearner
S
Sankaran Vaidyanathan
Basis Research Institute, University of Massachusetts Amherst
R
Rafal Urbaniak
Basis Research Institute
E
Emily Bunnapradist
Basis Research Institute
M
Michelangelo Naim
Basis Research Institute
Daniel Waxman
Daniel Waxman
PhD Candidate, Electrical Engineering, Stony Brook University
bayesian machine learninggaussian processesonline learningcausal inference