Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

📅 2026-08-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of assumption validity, robustness, and scalability in conditional independence (CI) testing for constraint-based causal discovery with high-dimensional mixed-type data. It systematically reviews and compares six major classes of CI methods—partial correlation, contingency tables, regression residuals, k-nearest neighbors, kernel-based approaches, and machine learning techniques—evaluating their performance and failure modes under diverse data-generating mechanisms, small sample sizes, and heterogeneous variable types. For the first time, it comprehensively delineates the applicability boundaries of each method and elucidates how their errors propagate into inaccuracies in causal skeleton and v-structure identification. The work also surveys current implementations in R and Python libraries and identifies key future directions, including discretization-free CI tests for mixed data, improved error control in small samples, and enhanced scalability.
📝 Abstract
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.
Problem

Research questions and friction points this paper is trying to address.

Conditional Independence
Causal Discovery
High-dimensional Data
Mixed-type Variables
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Independence Testing
Causal Discovery
Robustness
High-Dimensional Data
Mixed-Type Data
💼 Related Jobs
No related jobs found.
P
Pavel Averin
Department of Computer Science, School of Sciences and Engineering, University of Nicosia, 2417, Nicosia, Cyprus
T
Theodoros Moysiadis
Department of Computer Science, School of Sciences and Engineering, University of Nicosia, 2417, Nicosia, Cyprus
Ioannis Katakis
Ioannis Katakis
Department of Computer Science, School of Sciences and Engineering, University of Nicosia, 2417, Nicosia, Cyprus