Score
Designs and applies diagnostic procedures and statistical analyses to assess whether exchangeability assumptions hold between comparison groups or conditions. This work includes testing covariate balance, evaluating shared-environment and independence conditions, and quantifying the sensitivity of inferences to violations of those assumptions.
This paper addresses the practical violation of conditional unconfoundedness—a core assumption in causal machine learning—by systematically analyzing the bias mechanisms of causal forests and X-learners under unmeasured confounding. It introduces negative control outcomes (NCOs) as a novel, pragmatic diagnostic tool for detecting unobserved confounding, marking the first application of NCOs for this purpose in causal ML. Through extensive simulation studies across diverse confounding structures, sample sizes, and degrees of treatment effect heterogeneity, the study demonstrates that unconfoundedness violations induce spurious heterogeneity detection. Although NCOs do not strictly satisfy theoretical identification conditions, they robustly flag subgroups severely affected by unmeasured confounding, thereby substantially improving the credibility of individualized treatment effect estimation. This work advances the integration of NCOs from sensitivity analysis into standard causal modeling pipelines, providing methodological foundations for robust causal machine learning.
This paper addresses the often-overlooked exchangeability assumption underlying permutation tests in multiple linear regression—a condition critical for valid statistical inference. We rigorously clarify the logical relationship between exchangeability and the null hypothesis, and systematically evaluate the robustness of common permutation schemes—response permutation, residual permutation, and design-matrix permutation—under both satisfied and violated exchangeability conditions, via theoretical analysis and simulation studies. We propose a novel, pedagogically integrated framework that unifies conceptual understanding with formal theory, and extend the analysis for the first time to hierarchical and clustered regression models, enhancing methodological generality. Results show that while standard regression settings yield consistent conclusions across permutation schemes, inference deteriorates markedly when exchangeability fails. Our framework significantly improves students’ conceptual grasp of resampling-based inference, offering a new paradigm for statistics education and applied practice.
Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.
In causal inference, conventional covariate balance tests suffer from inflated false-positive rates when irrelevant covariates are imbalanced and exhibit low sensitivity to imbalance in potential outcomes. To address these limitations, we propose a conditional balance test grounded in prognostic covariate importance—explicitly incorporating each covariate’s predictive strength for potential outcomes into the test weighting scheme. This enables joint optimization of statistical power and false-positive control. Our method employs a standardized regression-weighted mean difference test, supported by theory-driven weight construction and a Monte Carlo simulation validation framework. We provide theoretical guarantees of improved statistical power. Simulation studies demonstrate that our approach achieves substantially higher detection power than global balance tests under potential outcome imbalance, while reducing the false rejection rate due to irrelevant covariate imbalance by over 40%.
In cross-population causal inference, unobserved differences between source and target populations frequently violate the conditional exchangeability assumption—an untestable condition critical for external validity. To address this, we propose SLOPE (Sensitivity to Local Policy Exogeneity), the first metric that integrates the Hampel-type derivative robustness framework with influence functions to quantify an estimator’s sensitivity to local violations of conditional exchangeability. SLOPE admits a closed-form analytical expression, enabling robustness comparisons across source/target population pairs or estimation strategies—even when the underlying assumptions are empirically unverifiable. We validate SLOPE through reanalyses of multi-country randomized experiments, demonstrating its ability to identify more robust generalization pathways. Our results enhance the reliability and reproducibility of causal inference designs in heterogeneous populations.
This study addresses the challenge of population distribution shift that arises when external control data are incorporated into randomized controlled trials due to cost constraints, which violates the conventional exchangeability assumption and biases causal effect estimation. To overcome this limitation, the authors propose a distribution-shift-aware semiparametric framework that explicitly models the distributional discrepancy between trial participants and external controls. By integrating calibration equations to adjust the efficient influence function and employing an adaptive shrinkage strategy, the method constructs an augmented estimator that maintains consistency while achieving higher statistical efficiency than estimators relying solely on trial data. Both theoretical analysis and empirical evaluations demonstrate that the proposed approach substantially improves estimation efficiency across synthetic and real-world scenarios, effectively relaxing the stringent exchangeability requirement.
In small-sample randomized clinical trials, inference for covariate-adjusted risk difference estimates lacks methods that simultaneously ensure robustness, efficiency, and proper Type I error control. This study investigates the performance of unconditional exact tests, the Mantel–Haenszel method, and several g-computation approaches—including standard, robust, and penalized variants—through simulation. The results reveal that Type I error inflation primarily stems from a mismatch between the target parameter, variance estimation, and inferential objective, rather than merely from limited sample size. Accordingly, the work proposes a principled criterion for method selection that aligns these components: standard g-computation often leads to inflated Type I error in very small samples, whereas its robust or penalized alternatives improve error control at the cost of reduced power; classical methods like Mantel–Haenszel, while conservative, demonstrate consistent robustness.
This work addresses the challenge that, under non-exchangeability, the covariance structure of permutation statistics deviates from that of the original test statistics, rendering conventional studentization incapable of recovering the correct joint asymptotic distribution. To overcome this limitation, the authors propose a general and computationally efficient covariance correction method that requires no assumptions about specific parameters, test statistics, or permutation schemes, and remains valid even in singular covariance settings. Unlike existing approaches—such as pre-pivoting—which suffer from high computational costs, the proposed method accurately restores the asymptotic dependence structure of permutation statistics. Theoretical analysis and extensive simulations demonstrate that it achieves asymptotically valid and powerful multiple testing across diverse scenarios, significantly outperforming current methods in inferential accuracy and efficiency.