Score
Designs and implements procedures that quantify how well a proposed high‑level causal abstraction matches a lower‑level causal system by computing a continuous error or validity score (causal abstraction error). This competence covers defining intervention sets and variable mappings, and building faithfulness tests and statistical procedures that reliably discriminate valid versus invalid abstractions and converge with few sampled interventions.
This work addresses the lack of a unified validation framework for evaluating whether high-level causal abstractions faithfully reflect underlying mechanisms. The authors construct a benchmark encompassing ten classes of complex systems—spanning discrete/continuous and static/dynamic types—and systematically assess over thirty metrics under a common causal abstraction framework to distinguish valid from invalid abstractions. They propose a novel continuous measure, Causal Abstraction Error (CAE), which passes discriminative tests across all systems and converges with only 30 interventions. Additionally, they introduce a fidelity test for unmapped variables that integrates observational, functional, information-theoretic, and causal criteria. Experiments demonstrate that causal metrics constrained solely by faithfulness reliably discriminate abstraction validity, with CAE exhibiting both superior performance and computational efficiency.
This paper addresses the identifiability of causal effects under causal abstraction—specifically, how to infer treatment effects when the underlying causal graph is incompletely known, as commonly encountered in high-dimensional or complex systems. Method: We propose the first hierarchical identifiability criterion system tailored to causal abstraction, systematically linking abstract causal structures at varying granularities to their corresponding identification capabilities. Our theoretical framework integrates causal graph models, observational data, and formal logical reasoning, yielding decidable, layered identifiability criteria. Contribution/Results: The framework does not require a fully specified causal graph; it enables identifiability assessment even when the causal structure is unknown or only partially known. We validate its effectiveness and practicality through rigorous analysis of canonical examples from the causal inference literature.
Existing causal abstraction frameworks struggle with lossy abstractions—i.e., many-to-one mappings—because the conventional “abstraction invariance” assumption fails when multiple low-level interventions map to a single high-level intervention. To address this, we propose the **projection abstraction framework**, which relaxes this assumption and establishes the first rigorous causal abstraction theory for lossy representations. Our approach constructs learnable projection mappings that preserve causal semantics when transforming from complex, low-dimensional models to concise, high-dimensional abstract models. We introduce a graph-structural identifiability criterion enabling high-level causal structure inference from finite observational data. Theoretically, we prove cross-level transferability of the framework for observational, interventional, and counterfactual queries. Empirical evaluation on high-dimensional image domains demonstrates accurate recovery of high-level causal structures, significantly enhancing both interpretability and computational efficiency in modeling complex systems. (149 words)
This work addresses the lack of rigorous reliability evaluation for causal probing interventions in large language models. We propose the first quantifiable and comparable two-dimensional empirical framework, formalizing intervention effectiveness via “completeness” and “selectivity,” and defining their harmonic mean as the core “reliability” metric. Through hierarchical, controlled intervention experiments and cross-method benchmarking, we formally uncover fundamental reliability trade-offs: no single method achieves universal reliability across all network layers; nonlinear interventions outperform linear ones in shallow-to-middle layers, whereas linear interventions exhibit greater robustness in deeper layers; and concept removal methods are significantly less reliable than counterfactual interventions—challenging their validity for causal explanation. Our framework establishes a theoretical benchmark and practical guidelines for causal interpretability research in foundation models.
This study systematically evaluates judgment biases of large language models (LLMs) in causal reasoning, revealing a “skepticism trap” at Level 1 of Pearl’s causal hierarchy—where models like Claude Haiku erroneously reject 60% of valid causal chains—and a “non-monotonic scaling paradox” at Level 3, exemplified by GPT-5.2 underperforming GPT-4-Turbo by 55 points. To address these issues, we introduce the T3 benchmark, grounded in Pearl’s ladder of causation and comprising 454 expert-crafted scenarios, assessing model performance across utility, safety, and calibrated refusal. We further propose a structured process verification protocol (RCA) and adversarial fuzzy counterfactual testing, complemented by high-resolution failure analysis, which collectively enhance both decisiveness and accuracy in causal judgments.
Traditional voting mechanisms in causal reasoning often fail due to answer dispersion or repetitive errors. This work proposes CALVER, a training-free symbolic verifier that introduces, for the first time, axiom-level causal validation into large language model (LLM) reasoning selection, enabling identification of graph-structurally valid causal paths without relying on reference answers. CALVER leverages Pearl’s causal criteria—including d-separation, backdoor adjustment, and interventions—combined with text-derived causal graphs and Bayesian networks to achieve millisecond-level scoring efficiency on CPU. Evaluated on the CLEAR benchmark, it attains an accuracy of 42.1%, substantially outperforming majority voting, reward models, LLM judges, and confidence-based baselines (all around 30%), and demonstrates consistent superiority across diverse models, network architectures, and causal graph construction scenarios.
Existing causal abstraction methods support only global evaluation of explanation faithfulness, making it difficult to diagnose where explanations succeed or fail across specific input regions. This work introduces input space partitioning into the causal abstraction framework, enabling fine-grained localization and attribution of explanation efficacy by identifying “well-explained” and “under-explained” regions through single-intervention swaps. The approach not only reveals failure modes of high-level causal hypotheses but also offers recursive reconstruction and compositional strategies to refine explanations. Experiments demonstrate precise error analysis across multiple causal abstraction settings and show that, in toy logical tasks, the method can recover correct high-level hypotheses from scratch, thereby validating its effectiveness in constructing more accurate and scalable mechanistic explanations.
In real-world scenarios where ground-truth causal relationships are unavailable, assessing the reliability of numerous pairwise causal statements remains challenging. This work proposes a novel paradigm for causal evaluation that does not rely on the faithfulness assumption. It introduces a compatibility score to measure the consistency between causal statements and observed data, complemented by an incompatibility score derived from global structural constraints of directed acyclic graphs, thereby jointly evaluating the overall plausibility of all binary causal relations. The approach integrates linear causal models, graphical model theory, mutual information analysis, and techniques for evaluating large language model outputs. Both theoretical analysis and empirical experiments demonstrate that the proposed scoring framework effectively discriminates between correct and erroneous causal claims and is successfully applied to assess causality assertions generated by large language models.
This work addresses a critical limitation in existing interventional interpretability evaluations, which rely on point estimates and struggle to disentangle true causal effects from sampling or adaptive biases. The authors reformulate the problem as a causal estimation task and introduce, for the first time, an anytime-valid statistical certification framework that accommodates adaptive intervention sampling. By integrating Hoeffding-type confidence sequences with variance-adaptive betting strategies and bounded mixture importance weighting, the method yields both confidence intervals and dynamic confidence sequences for intervention fidelity. Empirical validation on MNIST abstractions and GPT-2 Small IOI circuit experiments demonstrates that the approach not only certifies high-fidelity interpretability claims and detects statistically insignificant differences between methods but also reduces certification costs by 10–30× compared to existing baselines.