Score
Deriving conditions, parameterizations, and estimators that guarantee unique, consistent recovery of causal or structural quantities from observed data—e.g., sparse-safe parametrizations, mediation identification with multiple mediators, and identification under known nonlinearities.
This paper addresses the identifiability of causal effects in causal graphs incorporating additional logical or structural constraints—extending classical identifiability theory by introducing the novel notion of *constrained identifiability*. Method: We propose the first systematic decision framework based on arithmetic circuits (ACs), uniformly encoding causal graphs, parametric constraints, and logical conditions; we formally prove its completeness is at least as strong as do-calculus and that it rigorously handles implicit assumptions such as strict positivity. Results: Experiments demonstrate that incorporating logical or structural constraints renders several classically non-identifiable causal effects identifiable. Our work establishes a formal and computationally grounded foundation for knowledge-augmented causal inference, enabling principled integration of domain-specific prior knowledge into causal effect estimation.
This paper addresses the identifiability of total causal effects under multiple interventions in time-series settings where only a summary causal graph is known—specifically, whether such effects admit a do-free observational equivalent expression. We establish the first necessary and sufficient condition for adjustment-based identifiability under summary causal graphs and prove its theoretical completeness. Based on this, we design a pseudo-linear-time algorithm (nearly O(|E|)) for identifiability determination, substantially outperforming existing approaches. The algorithm uniformly handles multiple interventions, temporal dependencies, and sparse graph structures, enabling efficient causal effect identification in high-dimensional settings. Our main contributions are: (i) the first theoretical framework for adjustment criteria tailored to time-series summary graphs; (ii) a decidable, computationally tractable, and scalable do-free identification procedure; and (iii) rigorous guarantees of soundness and completeness for both the criterion and the algorithm.
This study addresses the problem of testing whether a treatment effect operates entirely through observed mediators and identifying causal mechanisms under control for covariates. The authors propose a statistical test based on double machine learning, extending— for the first time—the joint evaluation of full mediation and causal mechanism identification to non-randomized treatment settings. By integrating conditional independence testing, the method achieves root-n consistent and asymptotically normal inference even in the presence of high-dimensional covariates. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to two randomized experiments examining maternal mental health and social norms.
This work addresses the structural identifiability of parameters in stochastic differential equation (SDE) models under multiple interventions—i.e., whether SDE parameters can be uniquely recovered from samples of post-intervention stationary distributions. Theoretically, we establish the first uniqueness guarantee for SDE parameter recovery under multi-intervention settings; for linear SDEs, we derive a tight lower bound on the minimum number of required interventions; for weak-noise nonlinear SDEs, we obtain an upper bound on identifiability. Methodologically, we propose a parametric framework featuring learnable activation functions, integrating intervention modeling, stationary distribution analysis, and weak-noise asymptotic theory. Experiments on synthetic data demonstrate that our approach accurately recovers ground-truth parameters, and the theory-guided learnable architecture significantly improves both estimation accuracy and robustness.
Nonparametric causal mediation analysis with multivariate, continuous, and high-dimensional mediators remains challenging due to the lack of unified, efficient estimation frameworks. Method: We propose the first unified first-order bias-corrected estimation framework covering six major definitions of mediation effects. By reducing their nonparametric identification formulas to two fundamental statistical functionals, we develop a generic one-step estimator embeddable within arbitrary machine learning models (e.g., neural networks, random forests), achieving √n-consistency and asymptotic normality. Our approach integrates state-of-the-art semiparametric techniques—including targeted maximum likelihood estimation, Riesz representation learning, and density ratio estimation. Results: Simulations demonstrate substantial gains in estimation accuracy and robustness over existing methods. Applied to real clinical data, our method quantifies the mediating proportion of pain management in the pathway from chronic pain to opioid use disorder, confirming practical utility in healthcare analytics.
This study addresses the joint identification and counterfactual analysis in incomplete structural models featuring support and moment constraints. The authors embed counterfactuals directly into an augmented structural model, departing from the conventional “estimate-then-simulate” paradigm. By leveraging support function methods, they simultaneously achieve identification and inference, revealing a fundamental isomorphism between the two tasks. A key contribution is the formulation of irreducibility conditions that explicitly characterize all support implications. Under mild regularity assumptions, the support function approach preserves sharpness with respect to the moment closure—even in counterfactual settings where traditional sharpness fails. Moreover, for irreducible models, the identified set and the moment closure are statistically indistinguishable in finite samples.
Traditional instrumental variable methods rely on the stringent assumption that the structural equation model holds exactly—a condition often violated in practice, leading to invalid inference. This work proposes a novel inference framework based on debiased least squares and inverse problem regularization, which defines a target parameter that coincides with conventional estimands when the structural model is correctly specified yet remains well-defined and inferable even under model misspecification. By relaxing the requirement of exact structural equation validity, the approach ensures robust statistical inference under substantially weaker conditions, thereby significantly enhancing the reliability and applicability of instrumental variable methods in realistic settings.
This study addresses the problem of recovering the causal diffusion mechanism of a continuous-time sparse multivariate stochastic system from steady-state cross-sectional observations alone—a setting relevant to domains like gene expression where repeated temporal measurements are infeasible. Assuming the system follows a stationary diffusion process with a known causal graph structure, the authors propose a nonparametric kernel method to identify and consistently estimate its drift functions, thereby reconstructing the system’s infinitesimal time-evolution dynamics. This work establishes, for the first time, nonparametric identifiability of causal diffusion drift functions using only static data, without requiring time-series observations. By integrating cross-validation with low-frequency sampling theory, the method enables effective estimation. Theoretical analysis confirms the consistency of the estimator, and simulations demonstrate its empirical validity, opening a new avenue for inferring dynamic causal mechanisms from static snapshots.
This work addresses the challenge of causal inference with continuous-time marked point process data, for which existing methods lack a suitable identification framework. Building on martingale theory, the authors extend the core assumptions of discrete-time causal inference—consistency, exchangeability, and positivity—to the continuous-time setting. They formulate a dynamic treatment strategy and a potential outcomes model tailored to marked point processes and establish corresponding causal identification conditions. Leveraging this foundation, they derive a novel marginal g-formula that enables nonparametric identification of causal effects. The proposed framework subsumes existing results for discrete-time and counting process settings as special cases, demonstrating both theoretical compatibility and extensibility, thereby unifying survival analysis and causal inference within a coherent paradigm.
This work addresses the challenge of determining causal effect identifiability in linear structural causal models with latent confounding, a problem traditionally hindered by the double-exponential computational complexity of Gröbner basis methods. The authors propose a novel symbolic computation algorithm that, for the first time, decides rational identifiability of causal effects in quasipolynomial time and efficiently computes the lowest-degree identification formula under a given maximum degree constraint. By integrating techniques from algebraic geometry with causal inference theory, the method substantially enhances algorithmic scalability and practical applicability, thereby overcoming a longstanding computational bottleneck in the field.