Score
Designs and executes controlled interventions on model inputs, activations, or components—such as causal patching, ablations/knockouts, and corrupted-input patching—to test whether specific units, layers, attention heads, or probe modules causally contribute to model behavior. Builds diagnostic evaluation procedures that quantify the effects of interventions, compare potential versus actual mechanisms, and analyze outcomes to infer, validate, or reject hypothesized internal causal pathways.
Activation patching is widely used in mechanistic interpretability, yet its natural indirect effect (NIE) estimator conflates inter-component state-dependent interaction effects (INT), leading to misattribution of causal contributions. This work formalizes and identifies INT within activation patching through the lens of causal mediation analysis, demonstrating that such interactions are both unavoidable and decomposable. The authors propose reframing INT as a diagnostic tool for interpretability. By integrating compositional interaction decomposition with local affine analyses, they empirically show in GPT-2’s IOI circuit that INT can cause critical components to be overlooked or their importance inflated, thereby explaining the instability of faithfulness scores. Building on these insights, they introduce new criteria for prompt-dependence analysis and mechanism discovery grounded in INT.
Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.
This work addresses the testing challenge of high-dimensional, non-deterministic, and computationally expensive software systems—exemplified by the CARLA autonomous driving simulator—where latent variables and variable interactions undermine causal inference. Existing causal testing methods assume full observability and absence of interactions, rendering them inapplicable to realistic, partially observable settings. To overcome this, we introduce effect modification analysis and instrumental variable methods into software causal testing for the first time, establishing a robust verification framework capable of modeling latent variables and identifying interaction effects. Crucially, our approach requires neither full log recording nor source-code instrumentation; it achieves reliable validation of three system-level requirements in CARLA using only limited, controlled data under low observability. As a result, it substantially reduces dependence on large-scale test data and strong observability assumptions, advancing practical causal testing for complex cyber-physical systems.
This work addresses the lack of rigorous reliability evaluation for causal probing interventions in large language models. We propose the first quantifiable and comparable two-dimensional empirical framework, formalizing intervention effectiveness via “completeness” and “selectivity,” and defining their harmonic mean as the core “reliability” metric. Through hierarchical, controlled intervention experiments and cross-method benchmarking, we formally uncover fundamental reliability trade-offs: no single method achieves universal reliability across all network layers; nonlinear interventions outperform linear ones in shallow-to-middle layers, whereas linear interventions exhibit greater robustness in deeper layers; and concept removal methods are significantly less reliable than counterfactual interventions—challenging their validity for causal explanation. Our framework establishes a theoretical benchmark and practical guidelines for causal interpretability research in foundation models.
Current interpretability research for large language models (LLMs) treats interpretability and controllability as disjoint objectives. Method: This paper proposes “intervention capability” as a unified evaluation goal and introduces an encoder-decoder framework that integrates four method families—sparse autoencoders (SAEs), Logit Lens, Tuned Lens, and probes—to enable controllable interventions on interpretable features. Contribution/Results: We formally define two novel metrics—intervention success rate and consistency–intervention trade-off—and argue that effective intervention constitutes the foundational objective of interpretability. Experiments show that Lens-based methods outperform SAEs and probes in simple interventions; however, existing methods exhibit inconsistent cross-feature and cross-model intervention efficacy. Moreover, mechanistic interventions often underperform prompt engineering, revealing critical controllability bottlenecks. This work shifts LLM interpretability research from descriptive analysis toward causal, interventionist control.
This work proposes a “layered attribution” diagnostic framework to disentangle the origins of inscrutable behaviors exhibited by AI agents in complex social systems, which are often conflated between internal representations and external constraints. The framework systematically distinguishes a foundational computational layer—encompassing architecture, memory, and perception—from a behavioral modulation layer comprising identity, goals, social interactions, and institutional constraints, thereby integrating representation learning, multi-agent modeling, and institutional analysis into a unified two-tier diagnostic architecture. It yields three key insights: behavioral substitutability validity hinges on the coupling among model, task, and layer; human–AI behavioral discrepancies can serve as diagnostic signals; and effective governance presupposes precise source attribution. This approach establishes a theoretical foundation for interpreting and governing AI behavior.