Score
Designs and implements statistical models for longitudinal (panel) and time‑series data to estimate causal and structural relationships, including panel regressions, fixed‑effects specifications, and time‑series econometric models. Builds and cleans unit–time datasets, estimates parameters and treatment effects, conducts inference and diagnostic tests (e.g., for serial correlation and heteroskedasticity), handles unit entry/exit, and produces model‑based predictions and policy‑relevant estimates.
Conventional causal inference methods for panel data rely heavily on strong parametric assumptions—such as linearity and fixed effects—limiting robustness and generalizability. Method: This paper proposes a nonparametric, matching-based analytical framework specifically designed for time-series cross-sectional (TSCS) data. It introduces the first systematic dynamic matching approach for such data, integrating propensity score matching, covariate balance diagnostics, and interactive visualization (via ggplot2 and Shiny). Contribution/Results: The framework relaxes assumptions on functional form and individual homogeneity, enhances assumption robustness and result verifiability, and delivers comprehensive diagnostic reports alongside interpretability assessment tools. Implemented as an open-source R package, it has been empirically validated across sociology, economics, and medical research, improving transparency, reproducibility, and credibility of longitudinal causal effect estimation.
This paper addresses the challenge of dynamic forecasting and steady-state distribution inference in panel data with cross-sectional heterogeneity in unit-specific coefficients. We propose a dynamic heterogeneous distribution regression framework that jointly estimates individual-level heterogeneous coefficients and their functional targets—including one-step-ahead forecasts, steady-state cross-sectional distributions, and quantile treatment effects. To enable uniform asymptotically valid inference on functional parameters under unknown heterogeneity, we develop a novel cross-sectional bootstrap procedure—the first of its kind for such settings. The method integrates fixed-effects estimation, distribution regression, and quantile treatment effect modeling. Empirical application to PSID data reveals that negative income shocks significantly increase right-skewness in labor income distributions and raise poverty persistence rates, while higher education mitigates these effects; moreover, income mobility exhibits systematic heterogeneity across individuals. Simulation studies confirm the method’s robustness and reliability.
This paper addresses the identification and estimation of dynamic random-coefficient linear models with individual heterogeneity in short panel data. Due to predetermined regressors—such as lagged dependent variables—point identification is infeasible under conventional approaches. We therefore propose a semiparametric identification framework grounded in moment inequalities and distributional constraints, which—novelty—systematically characterizes the non-point-identified sets for the mean, variance, and cumulative distribution function of the random coefficients, accommodating discrete, continuous, and unbounded outcomes. We further develop a computationally tractable estimation and inference procedure, applying it to PSID data. Empirically, we find substantial unobserved heterogeneity in U.S. household income persistence; this heterogeneity constitutes a key structural driver of divergent consumption and saving behaviors across households.
Estimating causal effects of state-level opioid policies—such as Prescription Drug Monitoring Programs (PDMPs), naloxone access laws (NALs), and medical cannabis laws—is challenged by staggered implementation, small sample sizes, and dynamic policy evolution, rendering conventional difference-in-differences (DID) and synthetic control methods invalid. Method: This paper proposes a novel causal inference framework based on autoregressive models, specifically designed for staggered, multiple-policy settings. It introduces formal identification assumptions for autoregressive structures, rigorously characterizes bias sources, and conducts principled sensitivity analysis. Contribution/Results: The framework unifies theoretical interpretability with empirical robustness in complex, overlapping policy environments. Simulation studies and real-world policy evaluations demonstrate that the method significantly outperforms existing mainstream approaches in estimation accuracy, stability, and reliability of counterfactual predictions.
This paper addresses the failure of causal identification in longitudinal panel data due to spatiotemporal interference—where an individual’s outcome is affected by others’ past treatment assignments. We propose a design-based causal inference framework that, under minimal assumptions (unknown interference structure and sequential ignorability), formally defines and identifies separable direct effects and spatiotemporal spillover effects for the first time. We demonstrate that conventional fixed-effects and difference-in-differences (DID) estimators suffer from systematic bias under interference. To overcome this, we construct a new estimator with consistency and asymptotic normality. Theoretical analysis, Monte Carlo simulations, and replications of two canonical empirical studies validate our approach: it substantially reduces estimation bias in spillover effects and effectively corrects the failure of standard panel methods under complex interference patterns.
This study addresses efficient estimation and robust inference for semiparametric and nonparametric models with fixed effects in panel data. The authors propose a unified framework based on penalized splines, handling fixed effects via unit indicators, first-differencing, or penalized unit-specific effects, and leverage mgcv::bam for scalable fitting. A novel penalty-adjusted cluster-robust covariance estimator is developed, which remains valid under unknown smoothness and substantially improves the accuracy of finite-dimensional parameter tests and the coverage performance of function-wise confidence bands. Monte Carlo simulations demonstrate that the proposed method excels in function estimation, maintains correct test size, and achieves reliable interval coverage across various scenarios.
论文探讨了在难以使用非结构化协方差矩阵时,采用高阶自回归模型作为处理制药行业面板数据相关性的一种有效方法。
This study addresses the challenge of identifying causal effects in nonseparable models when unobserved time-varying individual heterogeneity is correlated with explanatory variables. The authors propose a novel approach that approximates the conditional average potential outcomes using linear sieve methods, combined with individual-specific ridge regression and bias correction. Within a large-T asymptotic framework, this method achieves point identification, circumventing the partial identification issues common in traditional approaches. The resulting estimator admits an empirical Bayes interpretation, accommodates discrete treatment variables, and provides a unified framework for estimating average causal effects, counterfactual consumer welfare measures, and individual tax elasticities. An empirical application to supermarket scanner data demonstrates the method’s effectiveness by precisely quantifying the average equivalent variation and deadweight loss induced by price increases.
This study addresses causal inference in panel data under policy interventions by moving beyond the conventional parallel trends assumption. It proposes a novel framework grounded in factor models, wherein treatment effects arise from structural changes in the exposure of treated units to latent common shocks—and can further accommodate shifts in the factor process itself. The approach is applicable to settings with either a single or multiple treated units and remains effective even when unit- and time-specific effects are not separately identifiable. By incorporating treatment-dependent factor structures and combining fixed-effect estimation with tailored inference strategies, the method yields confidence intervals whose empirical coverage closely matches nominal levels in simulations. Empirical applications to California’s tobacco control program and German reunification produce results consistent with synthetic control estimates while, for the first time, enabling formal statistical inference.
This study addresses the problem of prospective causal forecasting in panel data—specifically, predicting counterfactual outcomes for a target unit over future periods during which an intervention has not yet been implemented. To this end, the authors propose the Two-Way Synthetic Forecasting (TWSF) estimator, which uniquely integrates synthetic control methods with multivariate time series extrapolation. The approach leverages a low-rank latent factor model to capture both cross-sectional dependencies and temporal dynamics, combining unit-specific regressors, time factor modeling, and orthogonalization-based bias correction. Theoretical analysis establishes pointwise consistency, asymptotic normality, and finite-sample error bounds for multi-step-ahead predictions. Extensive simulations confirm the estimator’s empirical performance, and a real-world application demonstrates its utility in evaluating the causal impact of 2020 NFL stadium reopenings on public health outcomes.