extract causal subgraphs

Designs and implements methods to identify and extract subgraphs that causally influence a target outcome or a model prediction on graph-structured data. This includes algorithms to estimate and quantify the causal effect of candidate subgraphs (often via interventions or perturbations), to search or rank subgraphs by their causal impact, and to produce explanations attributing model decisions to specific graph components.

extractcausalsubgraphs

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Survey on Causal Discovery: Theory and Practice

May 17, 2023
AZ
Alessio Zanga
🏛️ University of Milano - Bicocca

This study addresses the challenge of identifying causal graphs and estimating causal effects from observational data. We propose the first unified analytical framework that horizontally integrates major causal discovery paradigms—including constraint-based methods (e.g., PC), score-based methods (e.g., GES), functional causal models (e.g., LiNGAM, ANM, CAM, NOTEARS), and neural causal learning—while rigorously characterizing their identifiability conditions and practical applicability boundaries. Our contribution comprises: (1) a comprehensive knowledge graph covering 12 algorithmic families, 8 open-source toolkits, and applications across six domains (e.g., healthcare, economics, ecology); (2) standardized benchmark datasets, reproducible evaluation protocols, and practitioner-oriented guidelines; and (3) paradigm-level unification, formal identification boundary analysis, and an end-to-end resource ecosystem for real-world causal discovery deployment.

Identifying causal effects using graphical modelsReviewing algorithms and tools for causal inferenceSurveying causal discovery methods from data

Treatment Effect Estimation for Graph-Structured Targets

Dec 29, 2024
SH
Shonosuke Harada
🏛️ Kyoto University

To address bias in causal effect estimation on graph-structured data arising from node selection bias (e.g., preferential sampling of high-centrality nodes), this paper proposes GraphTEE—a novel framework that formalizes graph-level causal inference as a subgraph-level counterfactual reasoning task, the first of its kind. Methodologically, GraphTEE integrates graph neural networks, propensity score weighting, and confounder identification, augmented by a structured regularization mechanism grounded in confounder-aware clustering. This design enables theoretically grounded bias mitigation. Extensive experiments on synthetic and semi-synthetic graph datasets demonstrate that GraphTEE significantly outperforms state-of-the-art methods: it reduces mean absolute error by 23.6% and bias estimation error by 31.4%. These results validate GraphTEE’s robustness and interpretability for causal inference under complex, interdependent graph structures.

Graph DataObservation BiasTreatment Effect Estimation

Post-selection inference for causal effects after causal discovery

May 10, 2024
TC
Ting-Hsuan Chang
🏛️ Columbia University | Zhejiang University

Direct effect estimation on a selected causal graph induces selection bias due to data reuse, invalidating confidence intervals. Method: We propose the first post-selection inference framework for fixed-population causal effect parameters, integrating resampling with graph-structure screening to depart from the conventional “select-then-infer” paradigm. Built upon the PC algorithm, our approach unifies conditional independence testing, Gaussian modeling, and joint estimation over multiple candidate graphs, and is modularly extensible to other causal discovery algorithms and distribution families. Contribution/Results: We establish asymptotic validity—specifically, asymptotically exact coverage—for confidence sets targeting the true causal effect. Empirical evaluations demonstrate that our method substantially improves reliability and robustness of causal inference under uncertainty, yielding well-calibrated confidence sets even after graph selection.

Addresses invalid confidence intervals from data reuseEnsures correct coverage for true causal effect parameterGeneralizes approach across discovery algorithms and distributions

Explaining The Behavior Of Black-Box Prediction Algorithms With Causal Learning

Jun 03, 2020
NS
Numair Sani
🏛️ Sani Analytics | Columbia University | Johns Hopkins University

Existing interpretability methods for black-box image classification models suffer from insufficient causal grounding and fail to distinguish genuine causal features from spurious correlations induced by unobserved confounders. Method: This paper proposes an explainability framework grounded in interventionist counterfactual causal reasoning. It constructs a causal graph model incorporating unobserved confounders, integrates structure learning with latent variable modeling to abstract raw pixels into high-level semantic features, and identifies true “difference-makers”—features whose counterfactual interventions alter predictions—under arbitrary unmeasured confounding. Contribution/Results: This work introduces the first systematic application of counterfactual causal explanation to black-box model auditing, enabling verifiable identification of causal drivers. Experiments on image classification tasks demonstrate significant improvements in causal feature identification accuracy, thereby supporting rigorous algorithmic attribution analysis and trustworthy model evaluation.

Complex Prediction AlgorithmsExplainable AIFeature Importance

This paper addresses the problem of identifying the set of direct causes (i.e., local causal structure) of a target variable from purely observational data in a single environment—without interventions or full DAG modeling. It introduces a lightweight data-generation assumption, strictly weaker than standard causal discovery premises, imposing minimal distributional constraints on non-target variables. For the first time, it systematically establishes multiple identifiability conditions under the no-intervention, single-environment setting. Leveraging structural constraint theory, the authors design two robust algorithms that integrate conditional independence testing with score-based optimization within a finite-sample estimation framework. Evaluated on benchmark and real-world datasets, the proposed methods significantly outperform baselines such as ICP—achieving higher accuracy and greater robustness. This work provides both theoretically more permissive and practically more viable foundations for local causal inference.

Developing algorithms to estimate direct causes without interventionsIdentifying direct causes of a target variable from observational dataRelaxing identifiability assumptions for local causal structure learning

Latest Papers

What's happening recently
View more

This work addresses a fundamental issue in graph representation learning: node aggregation operations violate core assumptions of causal inference, thereby compromising causal validity. To resolve this, the authors propose a causal modeling framework grounded in the graph’s minimal inseparable units, which rigorously ensures identifiability of the underlying causal structure. They further analyze the trade-offs and simplifying conditions required for exact causal modeling. Building on this theoretical foundation, they design a plug-in causal enhancement module compatible with existing graph learning pipelines. Experiments on controlled synthetic datasets validate the theoretical claims and demonstrate that the proposed approach substantially improves the model’s capacity for causal reasoning.

causal inferencecausal subgraphscausal validity

This study addresses the challenge of efficiently selecting observational variables for causal effect identification in partially specified causal models. Focusing on iterative identification within semi-Markovian models, this work leverages causal graph theory to establish necessary and sufficient graphical conditions under which previously unidentifiable queries become identifiable. Furthermore, it designs an efficient localization algorithm to optimize navigation through the model space and streamline the search for optimal observation strategies. By refining underlying identification mechanisms to precisely pinpoint critical observation opportunities, this research significantly enhances both the efficiency and accuracy of constructing identification strategies in causal inference.

causal inferenceiterative identificationmodel specification

This study addresses the challenges of causal inference in the endogenous formation of social networks—specifically unobserved confounding, reverse causality, equilibrium dependence, and sampling bias—by proposing a design-based nonparametric identification framework. Leveraging random variation in initial ties and repeated observations in panel network data, the approach treats nodes and their potential outcomes as non-stochastic, thereby circumventing conventional assumptions of random sampling and asymptotic approximations. An application to professional service firm data reveals a significant positive causal effect of indirect connections on tie formation, whereas the influence of node degree and local density is weak and statistically unstable. These findings underscore the method’s strength in handling the endogeneity and equilibrium complexity inherent in network formation processes.

causal inferenceendogenous networksreverse causality

This study addresses the unclear impact of simulator assumptions and the frequent conflation of identifiability with out-of-distribution generalization in supervised causal discovery. To this end, it constructs a multidimensional taxonomy encompassing encoders, structural decoders, and training paradigms to systematically analyze how data and simulator assumptions complement observational information. Furthermore, this work proposes a novel paradigm that aligns evaluation metrics with identifiable graph targets, revealing the risk of spurious precision arising from mechanistic limitations. Ultimately, it establishes an analytical framework for guiding method comparisons and identifies key open problems in transfer learning, test-time adaptation, and uncertainty quantification. Collectively, these contributions provide a rigorous theoretical foundation for evaluating causal discovery methods.

EvaluationGeneralizationIdentifiability

Hot Scholars

NR

Nidhi Rastogi

Assistant Professor, Rochester Institute of Technology, NY
CybersecurityArtificial IntelligenceAutonomous VehiclesGraph Analytics
NA

Nicholas Asher

CNRS Research Director (DRCE), member ANITI
Semanticspragmaticsdiscourse and dialogueNLP
HK

Harshad Khadilkar

Data Science Lead, Tata Group; Visiting Associate Professor, IIT Bombay
Multi-Agent AIControlOptimisationReinforcement Learning
ZL

Zhicong Lu

Assistant Professor, George Mason University
HCIsocial computinglive streamingcreativity support