Score
Quantifying and interpreting the magnitude of associations or treatment effects (including interactions and fixed effects), computing estimates with uncertainty, and attributing observed differences to coarse versus fine-grained factors.
This study addresses the heterogeneity of treatment effects in clinical research, where conventional subgroup analyses lack individual-level predictive power and purely machine learning–based approaches often lack statistical guarantees. To bridge this gap, the authors propose a two-stage hybrid workflow: first, using formal statistical hypothesis testing to confirm the presence of heterogeneous treatment effects, then constructing an individualized treatment strategy evaluated via cross-fitted doubly robust estimation under a Neyman–Pearson risk constraint. This framework integrates the interpretability of statistical inference with the predictive strength of machine learning, yielding a transparent, auditable, and statistically principled approach to heterogeneity. The method demonstrates efficacy in both simulation studies and the ACTG 175 HIV trial, and is accompanied by a practical implementation checklist along with guidance for alignment with regulatory-oriented heterogeneous treatment effect (HTE) assessment protocols.
This paper addresses the problem of identifying **interpretable, high-effect heterogeneous subgroups** from conditional average treatment effect (CATE) estimation. We propose a rule-set-based subgroup discovery framework—e.g., “treatment effect is significant if (X_1 > 0) and (X_2 < 1)”—that jointly optimizes subgroup size and effect magnitude, yielding a Pareto-optimal rule frontier; sample splitting ensures valid statistical inference. Our contributions are threefold: (i) explicit encoding of high-dimensional interaction effects into human-readable logical rules; (ii) the first application of multi-objective optimization to CATE subgroup identification, balancing interpretability, statistical significance, and representativeness; and (iii) theoretical guarantees of asymptotic unbiasedness and coverage for the derived rule sets. In extensive simulations and real-world policy evaluation datasets, our method substantially outperforms existing black-box subgroup detection approaches, achieving both strong statistical power and decision transparency.
This study addresses the challenge in applied microeconomics of effectively synthesizing empirical evidence, predicting effect sizes in new contexts, and correcting for publication bias. It proposes an integrated methodological framework that combines systematic literature review, covariate reweighting for extrapolation, and selection bias correction techniques—applicable even with as few as three prior studies. The approach innovates by offering a transparent and reproducible pipeline for out-of-sample effect prediction and, for the first time, quantifies the extent to which publication bias distorts average treatment effects. Empirical results demonstrate that bias-corrected average effects amount to only 12%–21% of naive unweighted averages, substantially improving predictive accuracy and enhancing the relevance of findings for policy design.
Causal inference from the fusion of experimental and observational data is often biased due to untestable assumptions—namely, external validity and ignorability. This paper introduces the first double machine learning (DML) framework that enables testable detection of violations of these assumptions. Our method jointly models both data sources via a residual-debiasing semiparametrically efficient estimator, yielding consistent estimation of treatment effects. We rigorously prove a “no-free-lunch” theorem, establishing that correct assumption identification is fundamentally necessary for consistency. Evaluated on multiple simulations and three real-world case studies, our approach significantly outperforms existing fusion methods in both estimation accuracy and robustness. Crucially, it provides diagnostic capability to detect assumption violations while maintaining theoretical guarantees. The framework thus offers both principled theoretical foundations and practical tools for causal extrapolation.
This paper addresses the challenge of subgroup identification for heterogeneous treatment effect (HTE) estimation in causal inference. We propose an interpretable, nested nonparametric subgroup partitioning tree framework that integrates honest splitting, debiased machine learning (DML), and an aggregation tree structure. To our knowledge, this is the first method ensuring subgroup nesting while enabling unbiased, asymptotically normal statistical inference for average treatment effects (ATEs) within each subgroup. Unlike existing approaches, it mitigates p-hacking risks and jointly optimizes subgroup granularity, interpretability, and statistical validity. Simulation studies demonstrate substantially improved power for detecting heterogeneity. Empirical analysis of maternal smoking’s impact on newborn birth weight reveals systematic HTE patterns driven by parental characteristics and delivery-related factors.
This paper addresses point estimation and uncertainty quantification for treatment effect paths—such as dynamic effects and event-study designs—in policy evaluation. To overcome the looseness of conventional uniform confidence bands, which ignore correlations among path estimators, we propose two data-driven feasible bound methods. Our novel framework jointly enforces average-effect coverage guarantees and path smoothness constraints, integrating post-selection inference, smooth regularization, Monte Carlo simulation, and robust point estimation. The resulting confidence bands are substantially narrower while maintaining valid coverage, especially under high estimator correlation; our point estimator also demonstrates superior performance across diverse simulation settings. The key contribution is the first systematic incorporation of smoothness priors into path inference, thereby unifying statistical rigor with economic interpretability.
Quantifying uncertainty in two-way fixed-effects estimation for networked data remains challenging due to violations of standard independence and homogeneity assumptions. To address this, we propose Branched Fixed Effects (BFE): a method that partitions the sample into statistically independent branches to enable unbiased causal inference and precise uncertainty quantification for treatment effects. BFE is the first to systematically integrate sample splitting into fixed-effects models under network dependence, thereby relaxing conventional assumptions—such as effect homogeneity or restrictive higher-order dependency structures—implicit in traditional standard error estimators. We develop an efficient, scalable algorithm supporting parallel branch extraction and estimation, enabling application to large-scale network data. Empirical evaluation on the Veneto firm-wage benchmark dataset (Italy) demonstrates that BFE substantially improves inferential robustness and interpretability. By ensuring replicability across research teams and facilitating transparent dissemination of credible estimates, BFE establishes a new paradigm for trustworthy causal inference in network settings.
This study addresses the common reliance on unrealistic assumptions about average treatment effects in experimental and observational research designs. It proposes a novel paradigm that shifts focus from directly positing average effects to modeling the full distribution of individual treatment effects, from which more plausible assumptions about average effects can be derived. By integrating distributional modeling with cross-disciplinary case studies, the approach demonstrates its validity and utility across diverse fields—including medicine, economics, and psychology—offering researchers a principled, heterogeneity-aware framework for specifying effect sizes grounded in empirical realism rather than idealized assumptions.
This study addresses the multiple testing problem in matched observational studies with a single intervention and multiple endpoints. We propose a robust method that jointly controls the false discovery rate (FDR) and quantifies unmeasured confounding bias. Our key innovation is the first integration of FDR control with formal sensitivity analysis, achieved via integer programming and a hierarchical screening strategy to efficiently compute sensitivity sets—i.e., subsets of hypotheses remaining significant under varying magnitudes of unmeasured confounding—enabling conservative estimation of the true positive rate (TPR). The method supports simultaneous inference across the entire hypothesis space, balancing statistical power and robustness. Simulation studies and an empirical application investigating long-term effects of childhood abuse demonstrate that our approach reliably identifies high-confidence endpoint subsets even under substantial hidden bias, substantially improving the reproducibility and interpretability of exploratory analyses.
This study addresses a critical limitation in existing causal inference methods, which predominantly focus on average treatment effects while neglecting the stochastic variability in individual responses. To capture this uncertainty, the paper introduces the variance of treatment effects (VTE) and the conditional variance of treatment effects (CVTE) as central measures of causal response heterogeneity. Under relatively weak assumptions allowing for unobserved confounding, the authors establish the identifiability of these variance measures and propose a consistent nonparametric kernel-based estimation framework. Theoretical analysis demonstrates the convergence properties of the proposed estimators, and experiments on synthetic and semi-simulated data show that the method achieves estimation accuracy that either matches or surpasses current baselines, thereby moving beyond the conventional mean-centric paradigm in causal inference.
This study addresses the challenge of causal inference bias arising from sample attrition in surveys and field experiments by introducing conformal inference into settings with missing data. It proposes a unified framework that integrates counterfactual modeling, weighting, and imputation strategies to estimate individual treatment effects without relying on strong, untestable assumptions commonly required by traditional methods. The approach yields prediction intervals that are both robust and precise, enabling valid comparisons of treatment effects across retained participants, those lost to follow-up, and the full sample. Simulation studies demonstrate that the proposed framework achieves higher coverage rates while producing narrower intervals compared to existing methods. Reanalyses of two empirical datasets further reveal heterogeneous treatment effects across distinct subpopulations, underscoring the method’s practical utility.