Score
Designs and implements causal-inference analyses that estimate the effect of adoption by constructing matched samples of adopters and non-adopters and applying difference-in-differences estimators; builds matching procedures, estimates DiD models, and evaluates robustness through alternative trend specifications and related diagnostics.
This paper addresses causal inference in multi-period, multi-group panel data settings. We propose a unified framework based on group-level matching. Its core innovation is the introduction of a generalized matching condition that embeds difference-in-differences (DID), synthetic control methods (SCM), and synthetic DID (SDID) into a single theoretical framework, revealing their intrinsic complementarity and equivalence under the parallel trends assumption. Through regret analysis, we formally characterize—for the first time—the applicability boundaries of DID and SCM. Moreover, we develop asymptotically efficient statistical inference procedures tailored to synthetic control estimation. Empirical applications demonstrate that our framework substantially improves the robustness and interpretability of policy effect estimates, offering a systematic solution for causal identification in complex observational settings.
Randomized controlled trials (RCTs) are often infeasible in software engineering, hindering rigorous causal assessment of tools, processes, or guidelines on development outcomes (e.g., efficiency, quality, user experience). Method: We propose a statistical causal inference methodology grounded in observational data, integrating the potential outcomes framework, propensity score matching, and difference-in-differences to systematically address confounding bias and selection bias. Contribution/Results: This work pioneers the systematic application of formal causal inference paradigms to requirements engineering and software practice research, tailoring analytical workflows and evaluation criteria to the characteristics of software engineering data. Empirical validation demonstrates that our approach substantially improves internal validity and reproducibility of causal conclusions in non-experimental settings. By enabling robust, evidence-based causal claims from real-world development data, it strengthens the empirical foundation for translating research findings into industrial practice.
This work proposes CausalSE, a novel framework that systematically integrates structural causal models (SCMs) with propensity score matching to rigorously identify the true causal effects of interventions—such as prompt engineering—on large language model code generation performance. Addressing a critical limitation in traditional software engineering empirical studies, which often rely on statistical associations vulnerable to confounding bias, this study introduces Pearl’s causal inference paradigm into the field. Empirical evaluation on the Galeras dataset reveals that while conventional association-based analyses suggest complex prompts improve performance, causal analysis under CausalSE finds no significant treatment effect, thereby exposing false-positive conclusions arising from unaccounted confounders. The paper further provides a reproducible methodology for causal inference in software engineering contexts.
This paper addresses causal inference in two-period panel data under the “no pure control group” setting: all units receive a strictly positive, heterogeneous continuous treatment in period two, rendering conventional difference-in-differences (DID) inapplicable due to the absence of untreated (zero-dose) units. Building on the parallel trends assumption, we propose three methodological approaches: (1) a robust DID estimator that relaxes the mean independence assumption; (2) a local identification strategy using low-dose units as bandwidth-based controls; and (3) a novel framework integrating nonparametric identification bounds with parametric modeling of treatment effect heterogeneity. Relative to Pierce & Schott (2016) and Enikolopov et al. (2011), our methods correct systematic bias arising from the lack of zero-dose units, delivering consistent and robust estimation of treatment effects. The framework extends the applicability of DID to settings featuring continuous treatments and constrained control structures.
This paper addresses how treatment-group selection threatens the parallel trends assumption in Difference-in-Differences (DiD) estimation—a critical yet under-characterized identification challenge. Method: We formally characterize the empirical content of this threat and derive necessary and sufficient conditions for parallel trends to hold under general selection mechanisms. We propose a “selection-driven bias decomposition framework” that systematically partitions DiD estimation bias into selection effects and time-varying heterogeneity effects, and develop operational benchmarking strategies—both with and without covariates—grounded in causal inference theory, selection modeling, and sensitivity analysis. Contribution/Results: Applied to the National Supported Work (NSW) experiment reanalysis, our approach quantifies and corrects selection bias, substantially improving the credibility of DiD estimates and the robustness of causal conclusions.
This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.
This study addresses the challenge in panel data causal inference that the identifying assumptions of difference-in-differences (DID), matching (M), and their hybrid (DIDM) are non-nested, leaving researchers without a principled basis for selecting a primary estimator. The paper proposes a unified selection framework grounded in the minimax-regret criterion and demonstrates that, under a broad class of loss functions, the estimand associated with DIDM achieves minimax-regret optimality. Consequently, it recommends DIDM as the headline estimator, with conventional DID and matching estimates serving as robustness bounds. Both theoretical analysis and empirical applications validate the efficacy of this approach, leading to a practical reporting guideline: prioritize DIDM while presenting DID and matching results as boundary checks.
This study addresses the challenges of causal inference in the endogenous formation of social networks—specifically unobserved confounding, reverse causality, equilibrium dependence, and sampling bias—by proposing a design-based nonparametric identification framework. Leveraging random variation in initial ties and repeated observations in panel network data, the approach treats nodes and their potential outcomes as non-stochastic, thereby circumventing conventional assumptions of random sampling and asymptotic approximations. An application to professional service firm data reveals a significant positive causal effect of indirect connections on tie formation, whereas the influence of node degree and local density is weak and statistically unstable. These findings underscore the method’s strength in handling the endogeneity and equilibrium complexity inherent in network formation processes.
This study addresses the vulnerability of conventional difference-in-differences (DID) estimators to bias under post-treatment shocks, which arises from their reliance on the parallel trends assumption. To overcome this limitation, the authors propose a novel inference approach that dispenses with this assumption by constructing a DID-specific predictor based on pre-treatment outcome dynamics and embedding it within a conformal inference framework. This method explicitly models potential post-treatment shocks and leverages pre-treatment information to impose identification constraints, thereby enabling robust causal inference even when parallel trends fail to hold. The proposed procedure substantially enhances the reliability and applicability of DID estimates in settings characterized by non-parallel trends.
This study addresses the limitation of prior research, which has only established correlations between code coverage and defect introduction without adequately controlling for confounding factors. For the first time in real-world JavaScript/TypeScript open-source projects, we treat code coverage as a continuous exposure variable, construct a causal directed acyclic graph to identify confounders, and employ generalized propensity scores combined with doubly robust regression to estimate both the average treatment effect and the dose–response relationship between coverage and defect introduction. Our findings reveal a nonlinear causal effect—such as threshold effects or diminishing marginal returns—providing the first empirical evidence grounded in causal inference to inform the optimization of testing strategies.