Score
Designs and implements statistical tests, metrics, and algorithms that detect and quantify pseudo-collinearity—near-linear or rank-deficient relationships between observed variables induced by latent confounders—by computing measures such as delta/δ-pseudo-collinearity and flagging variable pairs that exhibit this behavior. Uses those detections to guide causal-discovery procedures (for example by constraining edge orientations, prioritizing or modifying conditional-independence tests, and inferring hidden confounders) so as to distinguish apparent direct causation from confounding.
This work addresses the challenge of causal discovery in the presence of nonlinear mechanisms and latent confounding by proposing a novel approach grounded in the Minimum Description Length (MDL) principle. The method explicitly models latent confounding and accommodates nonlinear causal relationships by minimizing the Luckiness Normalized Maximum Likelihood (LNML) code length. It introduces the innovative notion of Δ-pseudocollinearity to identify dependency structures induced by latent variables and integrates this with a greedy search algorithm, termed PCG-CD, to construct a causal discovery framework that dispenses with assumptions of linearity or causal sufficiency. Empirical evaluations demonstrate that the proposed method accurately infers directed causal relationships and effectively detects latent confounding across both synthetic and real-world datasets.
Conventional rank-based tests for causal discovery with mixed (continuous + discrete) data fail under discretization, leading to inflated Type I error rates. Method: We propose a permutation-based rank test that establishes, for the first time, the permutation exchangeability of cross-covariance matrix rank tests under discretization, enabling asymptotically exact significance control in the presence of confounding discrete variables. Our approach integrates permutation testing, rank-statistic construction, and asymptotic distribution estimation—without requiring strong continuity assumptions on variables. Results: Experiments on synthetic and real-world data—including psychometric ordinal variables—demonstrate strict Type I error control and significantly higher statistical power than state-of-the-art methods. The method successfully enables causal structure learning in mixed-variable settings, advancing practical applicability of constraint-based causal discovery to discretized and ordinal data.
This paper addresses the problem of testing conditional independence among four observed variables given a latent variable in nonlinear or nonparametric models—a task for which classical tetrad constraints are invalid outside linear Gaussian settings. Methodologically, we propose the generalized tetrad constraint, the first extension of tetrad-based constraints to nonparametric models. Leveraging kernelized covariance operators, U-statistics, and asymptotic distribution theory, we construct a statistically falsifiable test framework with provably controlled Type I error rate (≤ 0.215). Experiments on simulated data and two real-world datasets—moral attitudes and intelligence test scores—demonstrate near-perfect statistical power (≈1) and substantially improved identifiability of latent causal structures. Our core contribution is the first tetrad-based conditional independence test for nonlinear latent variable models that simultaneously provides rigorous theoretical guarantees and strong empirical performance.
In large-scale observational studies, complete observation and control of confounders is often infeasible, leading conventional regression models (e.g., linear, logistic, Cox) to yield spurious associations. To address this, we propose a latent-variable-driven confounding structure model. Using both real-world and synthetic data simulations, we quantify how residual spurious association decays as the number of controlled confounders increases, and derive its closed-form mathematical expression. Our results demonstrate that even after adjusting for over 20 confounders, highly implausible causal hypotheses may still appear statistically “confirmed”; residual bias arising from unobserved latent confounders proves systematic and persistent. This work exposes a fundamental limitation of standard regression in causal inference and provides a computationally tractable theoretical framework—along with empirical benchmarks—for quantifying confounding bias. It thereby advances rigor, transparency, and caution in causal interpretation within survey-based research.
Direct effect estimation on a selected causal graph induces selection bias due to data reuse, invalidating confidence intervals. Method: We propose the first post-selection inference framework for fixed-population causal effect parameters, integrating resampling with graph-structure screening to depart from the conventional “select-then-infer” paradigm. Built upon the PC algorithm, our approach unifies conditional independence testing, Gaussian modeling, and joint estimation over multiple candidate graphs, and is modularly extensible to other causal discovery algorithms and distribution families. Contribution/Results: We establish asymptotic validity—specifically, asymptotically exact coverage—for confidence sets targeting the true causal effect. Empirical evaluations demonstrate that our method substantially improves reliability and robustness of causal inference under uncertainty, yielding well-calibrated confidence sets even after graph selection.
This study addresses the challenges of assumption validity, robustness, and scalability in conditional independence (CI) testing for constraint-based causal discovery with high-dimensional mixed-type data. It systematically reviews and compares six major classes of CI methods—partial correlation, contingency tables, regression residuals, k-nearest neighbors, kernel-based approaches, and machine learning techniques—evaluating their performance and failure modes under diverse data-generating mechanisms, small sample sizes, and heterogeneous variable types. For the first time, it comprehensively delineates the applicability boundaries of each method and elucidates how their errors propagate into inaccuracies in causal skeleton and v-structure identification. The work also surveys current implementations in R and Python libraries and identifies key future directions, including discretization-free CI tests for mixed data, improved error control in small samples, and enhanced scalability.
This study addresses the identifiability of causal structures in the presence of latent variables under location-scale noise models, which generalize beyond additive noise assumptions. The authors establish, for the first time, that acyclic directed mixed graphs (ADMGs) satisfying the bow-free condition are identifiable under such models, and further provide sufficient conditions for identifiability of causal directions even when the bow-free condition is violated. Building on this theoretical foundation, they propose a two-stage algorithm, LSNM-UV. Experimental results demonstrate that the proposed method significantly outperforms existing approaches based on additive noise models on heteroscedastic data, thereby validating both the correctness and practical advantage of the developed theory.
This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.