Score
Designing and implementing placebo (falsification) tests and robustness checks to probe causal claims and sensitivity to confounders, pre-trends, and sample selection. This includes constructing out-of-sample and sector-neutral checks and testing whether observed associations persist under alternative specifications or placebo interventions.
In observational studies, the untestable no-unmeasured-confounding assumption severely undermines causal inference reliability. To address this, we propose a counterfactual falsification framework leveraging heterogeneous data from multiple sources. Our method is the first to formalize “causal mechanism independence” as a testable implication of no unmeasured confounding. It models cross-environment dependencies among causal mechanisms and conducts a two-stage statistical test using high-power independence measures (e.g., HSIC or KCSD). Unlike randomized experiments, our approach requires no intervention and enables cross-environment failure detection. Extensive evaluations on synthetic and real-world datasets demonstrate that our method achieves significantly higher detection power for unmeasured confounding than existing falsification techniques, while maintaining strict control over false positive rates. This work establishes the first operational, empirically verifiable framework for confounding diagnosis in observational causal inference.
This paper addresses the validity of negative control variable (NCV) falsification tests in instrumental variable (IV) design. We identify a critical flaw: conventional NCV tests implicitly impose untestable functional-form restrictions—beyond standard exclusion and independence assumptions—leading to false rejection of valid IVs. To resolve this, we develop a unified theoretical framework that formally characterizes the identification conditions required for NCVs and establishes principled variable selection criteria. We further uncover and correct the implicit assumption bias inherent in existing NCV tests, proposing robust implementation guidelines. Methodologically, our approach integrates causal inference, IV identification theory, and conditional independence testing. The resulting framework substantially enhances the robustness, interpretability, and empirical reliability of IV analyses.
This study investigates whether the self-repair capability of frozen small code models in non-retrainable settings stems from repeated exposure to failed code or relies on external executable falsification feedback. To address this, we introduce a falsifiable methodology comprising feedback decomposition, content-controlled placebo design, matched-generation-budget control experiments, and executable auditing. We conduct large-scale evaluations on HumanEval+ and MBPP+ benchmarks using frozen models ranging from 0.5B to 1.5B parameters. Results show that blind resampling solves 18 more tasks than naive retrying; significant repair efficacy occurs only when feedback includes executable counterexamples, whereas pure instructions or content-irrelevant placebos yield no measurable improvement. These findings demonstrate that effective self-repair depends critically on external falsifying information rather than mere self-restatement.
This paper addresses the low statistical power and poor resolution of conventional placebo tests in synthetic control method (SCM) causal inference under small-sample settings—particularly when α = 0.05 and the number of donor units *N* is small. We propose a leave-two-out randomization inference framework that rigorously controls Type I error rates in finite samples while substantially improving test resolution and statistical power, even under stringent significance levels (α < 1/*N*). Unlike permutation or rank-based tests, our framework accommodates non-uniform treatment assignment and integrates formal sensitivity analysis for robust causal inference. Empirical results demonstrate that, under moderate effect sizes, the proposed method achieves lower actual Type I error rates and higher statistical power compared to standard approaches.
This paper addresses the lack of a unified framework for assessing systematic error (bias) across causal and descriptive inference. We propose the first cross-paradigm, generalizable bias risk assessment method, integrating modeling assumptions, data-generating mechanisms, and inferential objectives to cover high-risk settings—including randomized controlled trials (RCTs), nonprobability sampling, and statistical extrapolation—beyond traditional medical RCT constraints. Our approach combines qualitative bias mapping, assumption sensitivity analysis, and standardized reporting criteria, mandating explicit documentation of untestable assumptions and model uncertainty. The framework has been adopted as a mandatory reporting requirement by leading journals and funding agencies, thereby enhancing the reliability, interpretability, and external validity of research findings. (132 words)
This study addresses widespread misconceptions in the practical application of the synthetic control method, particularly concerning its reliance on covariates, claims of robustness, and prevailing model selection criteria—assertions often lacking empirical validation and potentially undermining causal inference reliability. Through rigorous theoretical analysis and extensive simulation experiments, the paper systematically evaluates these common misunderstandings and compares the performance of standard implementations against alternative approaches. The findings uncover critical pitfalls in current practices and, grounded in empirical evidence, offer concrete recommendations for more principled implementation and interpretation. By doing so, the work provides researchers with a practical guide to significantly enhance the quality and credibility of causal inferences derived from synthetic control methods.
This study addresses a critical limitation of the traditional synthetic control method: when treatment effects are weak, systematic bias can shift the center of confidence intervals, leading to misleading inferences. To remedy this, the authors propose a novel time placebo–guided approach that explicitly quantifies and corrects for such bias. By retrospectively assigning placebo intervention dates within the observed panel and refitting the synthetic control model at each, the method directly estimates the bias distribution under the null hypothesis. This enables the construction of nonparametric confidence intervals calibrated to maintain nominal coverage regardless of the true effect trajectory. The proposed procedure achieves stable, bias-corrected inference with fixed interval width, substantially enhancing the robustness of causal conclusions in synthetic control applications.
This study addresses the frequent under-identification and inadequate interpretation of outliers, non-replicable findings, and highly influential studies in meta-analyses, which often compromise the robustness of conclusions. It clarifies conceptual distinctions among these three types of problematic studies and proposes a systematic diagnostic framework that integrates robust statistical methods, graphical diagnostic tools, and advanced modeling techniques accounting for sampling variance dependencies. This approach enables more accurate detection of anomalous studies while leveraging visualization to facilitate interpretation of their potential sources. By synthesizing recent methodological advances, the work offers meta-analysts practical diagnostic strategies and cautious interpretive guidance, substantially enhancing the reliability and transparency of meta-analytic results.
This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.
This study addresses the severe inflation of Type I error that arises in conventional Mendelian randomization (MR) analyses when exposure and outcome variables are unknown and a preliminary association screening is performed prior to MR. To overcome this issue, the authors propose, for the first time, an MR testing framework robust to such pre-screening procedures. They introduce a novel causal test statistic with screening invariance, which integrates external summary statistics to rigorously control the Type I error rate under the null hypothesis of no causal effect while substantially improving statistical power. Extensive simulations demonstrate that the proposed method consistently outperforms standard MR approaches across diverse scenarios, offering both enhanced reliability and efficiency.