Score
Designs and estimates structural equation models and causal graphs that incorporate explicit constraints on allowable edges, parameter values, and variable groupings (e.g., forbidden-edge constraints and block structure). Builds and analyzes constrained causal-discovery procedures and parameter-learning methods that operate on fragmented, partially observed, or multimodal data while enforcing those domain or user-specified constraints.
This paper addresses exact identification of directed mixed graphs in simple structural causal models (SCMs) featuring both cycles and latent confounders. Due to cyclic dependencies, the graph skeleton cannot be recovered from observational data alone; moreover, latent variables invalidate standard conditional independence (CI) tests. To overcome these challenges, the authors propose a unified causal discovery framework that jointly leverages observational and interventional data. Grounded in do-separation and σ-separation, the framework integrates CI and do-CI tests and designs minimal intervention strategies. Theoretically, it establishes the first tight lower bound on the number of interventions required per experiment. Algorithmically, it introduces bounded and unbounded variants that fully recover the graph structure—under the assumption of no bidirected edges between neighbors—achieving logarithmic-factor-optimal time complexity. Empirical evaluations confirm both practical effectiveness and theoretical optimality.
This study addresses the challenge of identifying causal graphs and estimating causal effects from observational data. We propose the first unified analytical framework that horizontally integrates major causal discovery paradigms—including constraint-based methods (e.g., PC), score-based methods (e.g., GES), functional causal models (e.g., LiNGAM, ANM, CAM, NOTEARS), and neural causal learning—while rigorously characterizing their identifiability conditions and practical applicability boundaries. Our contribution comprises: (1) a comprehensive knowledge graph covering 12 algorithmic families, 8 open-source toolkits, and applications across six domains (e.g., healthcare, economics, ecology); (2) standardized benchmark datasets, reproducible evaluation protocols, and practitioner-oriented guidelines; and (3) paradigm-level unification, formal identification boundary analysis, and an end-to-end resource ecosystem for real-world causal discovery deployment.
This paper addresses the problem of identifying the set of direct causes (i.e., local causal structure) of a target variable from purely observational data in a single environment—without interventions or full DAG modeling. It introduces a lightweight data-generation assumption, strictly weaker than standard causal discovery premises, imposing minimal distributional constraints on non-target variables. For the first time, it systematically establishes multiple identifiability conditions under the no-intervention, single-environment setting. Leveraging structural constraint theory, the authors design two robust algorithms that integrate conditional independence testing with score-based optimization within a finite-sample estimation framework. Evaluated on benchmark and real-world datasets, the proposed methods significantly outperform baselines such as ICP—achieving higher accuracy and greater robustness. This work provides both theoretically more permissive and practically more viable foundations for local causal inference.
Direct effect estimation on a selected causal graph induces selection bias due to data reuse, invalidating confidence intervals. Method: We propose the first post-selection inference framework for fixed-population causal effect parameters, integrating resampling with graph-structure screening to depart from the conventional “select-then-infer” paradigm. Built upon the PC algorithm, our approach unifies conditional independence testing, Gaussian modeling, and joint estimation over multiple candidate graphs, and is modularly extensible to other causal discovery algorithms and distribution families. Contribution/Results: We establish asymptotic validity—specifically, asymptotically exact coverage—for confidence sets targeting the true causal effect. Empirical evaluations demonstrate that our method substantially improves reliability and robustness of causal inference under uncertainty, yielding well-calibrated confidence sets even after graph selection.
This paper addresses the nonparametric identification of total causal effects in dynamic systems characterized by summary causal graphs—structures that may contain latent confounders, lack temporal annotations, and admit directed cycles. Method: We generalize the front-door criterion to such incompletely specified graphs, introducing novel graphical conditions: “summary front-door paths” and “blocking-type backdoor shielding.” Leveraging do-calculus and graph-theoretic identifiability theory, we derive a verifiable sufficient criterion for effect identification. Contribution/Results: Unlike conventional approaches reliant on valid adjustment sets, our criterion ensures identifiability even when no admissible adjustment set exists. It establishes the first rigorous, testable framework for causal inference on summary causal graphs, thereby substantially broadening the scope of both causal discovery and effect estimation in dynamic systems with complex, partially observed structures.
This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.
A persistent methodological divide exists between the Neyman–Rubin potential outcomes framework and graphical causal models (e.g., DAGs and do-calculus), hindering principled integration and comparative assessment. Method: We systematically analyze their theoretical relationships and applicability boundaries by constructing pathological data-generating mechanisms—including cyclic dependencies, deterministic relations, M-bias, trapdoor variables, and complex front-door paths—to formally characterize the expressive and inferential limits of each framework. Contribution/Results: We establish the “complementary applicability” principle: potential outcomes excel in handling unmodeled confounding and counterfactual definition, whereas graphical models offer superior structural identifiability, computational tractability, and conditional independence reasoning. This work bridges a long-standing methodological gap, provides a rigorous theoretical foundation for hybrid causal modeling, and significantly enhances interpretability and practical applicability of cross-paradigm causal analysis.
This paper addresses the problem of learning directed acyclic graphs (DAGs) from data generated by nonlinear additive noise models (ANMs) with Gaussian noise. We propose a convex mixed-integer programming method based on basis function expansion and group ℓ₀ regularization, enabling explicit control of edge sparsity and seamless integration of structural prior knowledge. Theoretically, we establish statistical consistency and optimization convergence guarantees, and support early stopping as well as verifiably optimal solutions. Leveraging maximum likelihood estimation, branch-and-bound, and optimality gap analysis, we derive tight statistical error bounds. Experiments demonstrate that our approach significantly outperforms state-of-the-art DAG learning algorithms on both synthetic and real-world high-dimensional datasets. To the best of our knowledge, this is the first method to achieve consistent graph structure recovery under nonlinear ANMs while providing verifiable optimality within a user-specified precision.
This paper investigates the testable implications of exclusion restrictions and shape constraints within the potential outcomes framework. Method: We propose the first general graphical-model-based framework that sharply and constructively characterizes all observable implications of generalized support-set constraints. Our approach innovatively encodes the support sets of potential response functions via graph structures, integrating convex geometric analysis, counterfactual identification theory, and semiparametric testing techniques to uniformly derive complete testable conditions across diverse causal settings—including instrumental variables, mediation, and interference. Contribution/Results: Unlike prior case-specific analyses, our framework enables systematic and scalable characterization of testability. Empirically, we apply it to the US Lung Health Study, successfully identifying spousal spillover effects, exposure mapping structures, and the persistence of temporal treatment effects—demonstrating both statistical power and real-world applicability.
This work addresses the challenge posed by latent variables, which can induce non-directed acyclic and non-unique causal graphs among observed variables, thereby undermining conventional causal invariance methods. The paper characterizes the structure of such latent-induced observational graphs and, for the first time, establishes rigorous necessary and sufficient conditions for causal invariance in settings involving latent confounders, explicitly delineating its applicability even when the observed graph is not a DAG. Under a multivariate Gaussian assumption, the authors develop a verifiable theoretical framework that integrates causal graphical models with hypothesis testing. When the derived conditions hold, this framework accurately identifies observable causal parents, substantially enhancing the reliability of causal discovery in the presence of latent interference.