Score
Designs and builds causal graph structures and associated structural priors or skeletons, including algorithms that assemble or infer directed relationships from heterogeneous inputs. Develops frameworks to fuse, transfer, and reconfigure causal graphs—such as combining multiple causal sources or applying interventions to modify graph structure—so the resulting causal constructions can be analyzed or reused.
This study addresses the challenge of identifying causal graphs and estimating causal effects from observational data. We propose the first unified analytical framework that horizontally integrates major causal discovery paradigms—including constraint-based methods (e.g., PC), score-based methods (e.g., GES), functional causal models (e.g., LiNGAM, ANM, CAM, NOTEARS), and neural causal learning—while rigorously characterizing their identifiability conditions and practical applicability boundaries. Our contribution comprises: (1) a comprehensive knowledge graph covering 12 algorithmic families, 8 open-source toolkits, and applications across six domains (e.g., healthcare, economics, ecology); (2) standardized benchmark datasets, reproducible evaluation protocols, and practitioner-oriented guidelines; and (3) paradigm-level unification, formal identification boundary analysis, and an end-to-end resource ecosystem for real-world causal discovery deployment.
This paper identifies severe structural instability in causal graphs within software engineering (SE): identical SE data yield substantially different causal graphs—often with contradictory causal conclusions—when processed by different causal discovery algorithms or subjected to minor perturbations. Method: We systematically evaluate four representative algorithms—PC, FCI, GES, and LiNGAM—across 23 SE datasets, quantifying structural robustness via Jaccard similarity of edge sets. We conduct rigorous robustness experiments: cross-project, cross-version, bootstrap resampling, and parameter perturbation. Contribution/Results: Over 50% of edges vary across generation conditions; causal inferences for three canonical SE tasks—configuration selection, project management, and defect prediction—exhibit high inconsistency. This work provides the first systematic quantification of causal graph fragility in SE, establishes the necessity of “causal graph robustness testing,” and challenges prevailing practices in SE causal inference.
This study investigates whether Bayesian networks can be mapped to probabilistic structural causal models (SCMs) and analyzes the implications of such a mapping for network structure and joint distributions. By introducing independent latent random variables, deterministic structural equations are extended into probabilistic form, establishing correspondences between the two frameworks at semantic, structural, and distributional levels. Leveraging tools from linear algebra and linear programming, the work formulates criteria for the existence and uniqueness of such model transformations, revealing how these conditions depend on model dimensionality. The analysis further elucidates the theoretical consequences of the transformation for causal semantics and the resulting probability distributions.
This paper addresses exact identification of directed mixed graphs in simple structural causal models (SCMs) featuring both cycles and latent confounders. Due to cyclic dependencies, the graph skeleton cannot be recovered from observational data alone; moreover, latent variables invalidate standard conditional independence (CI) tests. To overcome these challenges, the authors propose a unified causal discovery framework that jointly leverages observational and interventional data. Grounded in do-separation and σ-separation, the framework integrates CI and do-CI tests and designs minimal intervention strategies. Theoretically, it establishes the first tight lower bound on the number of interventions required per experiment. Algorithmically, it introduces bounded and unbounded variants that fully recover the graph structure—under the assumption of no bidirected edges between neighbors—achieving logarithmic-factor-optimal time complexity. Empirical evaluations confirm both practical effectiveness and theoretical optimality.
该研究通过图手术和do-算子在确定性无环结构因果模型中建立了精确对应,解决了两者操作等价性的数学表述问题。
This work addresses a fundamental issue in graph representation learning: node aggregation operations violate core assumptions of causal inference, thereby compromising causal validity. To resolve this, the authors propose a causal modeling framework grounded in the graph’s minimal inseparable units, which rigorously ensures identifiability of the underlying causal structure. They further analyze the trade-offs and simplifying conditions required for exact causal modeling. Building on this theoretical foundation, they design a plug-in causal enhancement module compatible with existing graph learning pipelines. Experiments on controlled synthetic datasets validate the theoretical claims and demonstrate that the proposed approach substantially improves the model’s capacity for causal reasoning.
This study addresses the unclear impact of simulator assumptions and the frequent conflation of identifiability with out-of-distribution generalization in supervised causal discovery. To this end, it constructs a multidimensional taxonomy encompassing encoders, structural decoders, and training paradigms to systematically analyze how data and simulator assumptions complement observational information. Furthermore, this work proposes a novel paradigm that aligns evaluation metrics with identifiable graph targets, revealing the risk of spurious precision arising from mechanistic limitations. Ultimately, it establishes an analytical framework for guiding method comparisons and identifies key open problems in transfer learning, test-time adaptation, and uncertainty quantification. Collectively, these contributions provide a rigorous theoretical foundation for evaluating causal discovery methods.
本文针对现有因果推理框架无法处理循环因果依赖的问题,提出了一种新的二部图因果模型(BGCMs),通过明确指定干预方程、目标变量及其值来解决标准干预的模糊性。
This work addresses the challenges of high computational complexity and structural identifiability in large-scale causal discovery. It proposes a novel framework that dynamically integrates expert background knowledge into the core search process of scalable causal discovery—rather than merely using it for post-processing—for the first time. By combining constraint-based reasoning with efficient graph learning algorithms, the method actively guides structural exploration during the search, substantially reducing the hypothesis space. Experimental results demonstrate that the framework not only lowers computational overhead but also significantly improves the accuracy and identifiability of the inferred causal graphs, making it particularly well-suited for large-scale scenarios where only partial causal structures need to be recovered.
This work addresses the challenge of providing causal explanations for rare events (outliers) by formally defining causal paths and establishing their testable conditions. Building upon structural equation models, it uniquely integrates causal paths with the theory of causal abstraction, enabling verification that relies solely on pathways relevant to the rare event rather than requiring a complete causal graph. This approach bridges the gap between intuitive causal explanations and rigorous modeling, thereby constructing a causal abstraction framework tailored to rare events and supporting formal validation of root cause analysis results.
本文提出CausalArena,通过统一协议和多种SCM类型解决因果发现评估中的多样性和预训练-评估重叠问题。