Score
A counterfactual causal-inference framework that defines causal effects via potential outcomes, used to identify, estimate, and reason about causal effects, mediation, and finite-sample properties under explicit assumptions.
Existing causal inference frameworks—such as the potential outcomes model—rely on untestable counterfactual assumptions and abstract probability distributions, hindering verifiable, population-specific decision support. This paper proposes a novel paradigm: “finite-population treatment effect prediction modeling,” which treats the target population as the fundamental modeling unit and recasts causal inference as an empirically falsifiable prediction problem. It abandons unverifiable independence assumptions, rigorously distinguishes statistical from scientific inference, and introduces three core methodological innovations: analyzable treatment assignment mechanisms, error-source diagnostics, and fully testable causal modeling. For the first time, this framework enables systematic empirical validation of causal assumptions, exposes the fundamental dependence of causal conclusions on model specification, and substantially enhances transparency, reproducibility, and policy relevance of causal reasoning.
This study addresses the lack of formal definitions in existing root cause analysis methods, which are often limited to root nodes in causal graphs or biased toward proximate causes. Within the potential outcomes framework, this work proposes the first counterfactual definition of root cause at the individual level and introduces a probabilistic measure—Probability of Root Condition (PRC)—to quantify the likelihood that a candidate set of variables constitutes a root cause for a specific outcome. Under standard causal assumptions, the authors derive an explicit identification formula for PRC by integrating causal mediation analysis with counterfactual reasoning, thereby establishing its identifiability. The effectiveness and practical utility of the proposed approach are demonstrated through two numerical examples, filling a critical gap in the formal theory of root cause analysis.
This paper addresses counterfactual estimation under unobserved confounding in data-rich settings. Methodologically, it proposes a unified framework integrating structural causal models (SCMs) with latent factor models (LFMs), formally bridging graphical models and the potential outcomes paradigm. It establishes identifiability conditions for the average treatment effect (ATE), average treatment effect on the treated (ATT), and average treatment effect on the untreated (ATU), and derives general consistency conditions for estimation via principal component regression (PCR), latent factor modeling, and nonparametric smoothness analysis. The key contribution is a theoretical proof—under mild smoothness assumptions—that PCR consistently estimates all three average treatment effects, substantially relaxing conventional requirements of linearity, low dimensionality, and strong functional-form restrictions. This yields a robust, scalable solution for causal inference in high-dimensional observational data with large sample sizes.
Core assumptions in causal discovery and inference—such as causal sufficiency, faithfulness, and the Markov condition—are inconsistently formalized and ambiguously operationalized across methodological traditions, hindering rigorous method selection under non-ideal conditions (e.g., observational constraints, limited domain knowledge). Method: We propose the first cross-framework unification framework, employing conceptual analysis, comparative modeling, and structured meta-review to systematically map how distinct paradigms model these assumptions and align them with practical inferential goals. This yields a decision guide and reusable comparison toolkit spanning the entire causal analysis lifecycle—from problem formulation to result interpretation. Contribution: Our work achieves the first cross-paradigm integration of the formal semantics and operational logic of causal assumptions. It significantly enhances the rigor and efficiency of causal methodology selection and design, particularly when background knowledge is incomplete or data are observationally constrained.
This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.
While existing large language model–based social simulations can generate realistic interactions, they lack causal semantics, limiting their ability to support reliable causal inference for governance interventions. This work introduces, for the first time, a systematic integration of necessity and sufficiency–based causal concepts into social simulation, establishing a counterfactual framework tailored for policy evaluation. The framework explicitly articulates the relationship between simulator fidelity and policy relevance, offering theoretical guidance for simulator design and defining the fidelity criteria necessary to support valid policy inferences. By doing so, it advances social simulation beyond mere plausibility toward genuine decision-support capability.
This study addresses the problem of testing whether a treatment effect operates entirely through observed mediators and identifying causal mechanisms under control for covariates. The authors propose a statistical test based on double machine learning, extending— for the first time—the joint evaluation of full mediation and causal mechanism identification to non-randomized treatment settings. By integrating conditional independence testing, the method achieves root-n consistent and asymptotically normal inference even in the presence of high-dimensional covariates. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to two randomized experiments examining maternal mental health and social norms.
This paper addresses the testability of unmeasured confounding in observational studies, aiming to determine whether valid causal inference is feasible. We propose the first statistically rigorous method to test the “no unmeasured confounding” assumption, achieved by formally establishing a mathematical correspondence between the potential outcomes framework and causal graph models—thereby clarifying the fundamental distinction between causal identification and conventional association-based inference. Our approach operates within linear structural equation models and leverages joint analysis of randomized controlled trial (RCT) data and observational data to calibrate statistical power and rigorously control Type I error. The key contribution is the first empirically implementable, reproducible diagnostic test for unmeasured confounding, providing practitioners with a practical tool to assess the credibility and scope of causal conclusions drawn from observational studies.
This study addresses the problem of retrospective counterfactual prediction: estimating the expected potential outcome for an individual under a different intervention, given their observed covariates and realized outcome. To this end, the authors propose a unified framework based on a cross-world correlation parameter ρ(x), which links observable and unobservable potential outcomes within the Neyman–Rubin superpopulation model, thereby moving beyond the restrictive assumptions of ρ = 0 or ρ = 1 commonly adopted in existing methods. By incorporating correlation-aware identification strategies, they construct asymptotically efficient estimators and valid prediction intervals. Theoretical analysis and empirical experiments demonstrate that the proposed approach achieves nominal coverage under standard causal assumptions and substantially outperforms current baselines.
This study addresses the fundamental challenge in causal inference that individual counterfactual outcomes are unobservable, which forces existing methods to rely on strong, often unverifiable assumptions. To overcome this limitation, the authors propose the Digital Twin Counterfactual (DTCF) framework, which constructs individual-level digital twins to simulate counterfactual responses. The framework introduces a novel five-tier hierarchical validation architecture that translates simulation fidelity into testable assumptions on observable data. It formally distinguishes between marginally verifiable causal estimands—such as average treatment effects (ATE), conditional average treatment effects (CATE), and quantile treatment effects (QTE)—and those requiring joint structural assumptions, such as the full distribution of individual treatment effects (ITE) or benefit probabilities. By integrating copula-based modeling, sensitivity analysis, and uncertainty quantification, DTCF enables verifiable estimation of marginal effects while explicitly delineating the assumptions and quantification tools necessary for joint-structure-dependent quantities.