Score
Statistical and causal inference methods for longitudinal cross-sectional (panel) datasets, including fixed/random effects, instrumental variables, and decomposition approaches to recover persistence, capacity, and causal impacts. Applied skills include constructing panel datasets, addressing endogeneity, and estimating effects like credit capacity or policy impacts across units and time.
Conventional causal inference methods for panel data rely heavily on strong parametric assumptions—such as linearity and fixed effects—limiting robustness and generalizability. Method: This paper proposes a nonparametric, matching-based analytical framework specifically designed for time-series cross-sectional (TSCS) data. It introduces the first systematic dynamic matching approach for such data, integrating propensity score matching, covariate balance diagnostics, and interactive visualization (via ggplot2 and Shiny). Contribution/Results: The framework relaxes assumptions on functional form and individual homogeneity, enhances assumption robustness and result verifiability, and delivers comprehensive diagnostic reports alongside interpretability assessment tools. Implemented as an open-source R package, it has been empirically validated across sociology, economics, and medical research, improving transparency, reproducibility, and credibility of longitudinal causal effect estimation.
This paper addresses the failure of causal identification in longitudinal panel data due to spatiotemporal interference—where an individual’s outcome is affected by others’ past treatment assignments. We propose a design-based causal inference framework that, under minimal assumptions (unknown interference structure and sequential ignorability), formally defines and identifies separable direct effects and spatiotemporal spillover effects for the first time. We demonstrate that conventional fixed-effects and difference-in-differences (DID) estimators suffer from systematic bias under interference. To overcome this, we construct a new estimator with consistency and asymptotic normality. Theoretical analysis, Monte Carlo simulations, and replications of two canonical empirical studies validate our approach: it substantially reduces estimation bias in spillover effects and effectively corrects the failure of standard panel methods under complex interference patterns.
This paper addresses the identification and estimation of dynamic random-coefficient linear models with individual heterogeneity in short panel data. Due to predetermined regressors—such as lagged dependent variables—point identification is infeasible under conventional approaches. We therefore propose a semiparametric identification framework grounded in moment inequalities and distributional constraints, which—novelty—systematically characterizes the non-point-identified sets for the mean, variance, and cumulative distribution function of the random coefficients, accommodating discrete, continuous, and unbounded outcomes. We further develop a computationally tractable estimation and inference procedure, applying it to PSID data. Empirically, we find substantial unobserved heterogeneity in U.S. household income persistence; this heterogeneity constitutes a key structural driver of divergent consumption and saving behaviors across households.
This paper addresses fixed-effect estimation in linear panel data models by proposing a distribution-free adaptive shrinkage estimator. The method minimizes mean squared error within a broad class of shrinkage estimators—achieving, for the first time, shrinkage optimality without distributional assumptions. It accommodates time-varying fixed effects and arbitrary serial dependence structures, while adaptively shrinking estimates by jointly modeling cross-sectional and temporal correlations. The estimator admits a closed-form expression and is computationally efficient, supporting one-period-ahead forecasting. Monte Carlo simulations and empirical applications demonstrate substantial improvements in noise reduction and predictive accuracy over conventional shrinkage approaches—particularly under weak distributional assumptions or strong serial correlation.
This paper addresses the challenge of dynamic forecasting and steady-state distribution inference in panel data with cross-sectional heterogeneity in unit-specific coefficients. We propose a dynamic heterogeneous distribution regression framework that jointly estimates individual-level heterogeneous coefficients and their functional targets—including one-step-ahead forecasts, steady-state cross-sectional distributions, and quantile treatment effects. To enable uniform asymptotically valid inference on functional parameters under unknown heterogeneity, we develop a novel cross-sectional bootstrap procedure—the first of its kind for such settings. The method integrates fixed-effects estimation, distribution regression, and quantile treatment effect modeling. Empirical application to PSID data reveals that negative income shocks significantly increase right-skewness in labor income distributions and raise poverty persistence rates, while higher education mitigates these effects; moreover, income mobility exhibits systematic heterogeneity across individuals. Simulation studies confirm the method’s robustness and reliability.
This study addresses the challenge of causal effect estimation under non-random missingness—such as that induced by treatment selection—by integrating matrix completion into the unified “partition–intervene–reconcile” framework of causal inference, thereby establishing theoretical connections with difference-in-differences and synthetic control methods. The proposed approach effectively mitigates the “last-mile problem” in causal inference when treatment observations are sparse and missingness is non-ignorable. Empirical analysis of the impact of right-to-carry gun laws on violent crime demonstrates the method’s feasibility and robustness in policy evaluation, offering a theoretically rigorous and practically valuable pathway for causal inference with panel data.
This study addresses causal inference in panel data under policy interventions by moving beyond the conventional parallel trends assumption. It proposes a novel framework grounded in factor models, wherein treatment effects arise from structural changes in the exposure of treated units to latent common shocks—and can further accommodate shifts in the factor process itself. The approach is applicable to settings with either a single or multiple treated units and remains effective even when unit- and time-specific effects are not separately identifiable. By incorporating treatment-dependent factor structures and combining fixed-effect estimation with tailored inference strategies, the method yields confidence intervals whose empirical coverage closely matches nominal levels in simulations. Empirical applications to California’s tobacco control program and German reunification produce results consistent with synthetic control estimates while, for the first time, enabling formal statistical inference.
This study addresses the challenge of identifying causal effects in static panel data when the treatment is endogenous and high-dimensional nonlinear confounders are present, a setting where conventional instrumental variable (IV) methods often fail. The paper proposes Panel IV-DML, the first extension of double machine learning (DML) to a panel IV framework, which integrates flexible machine learning techniques—such as Lasso and random forests—for covariate adjustment and introduces a novel weak identification diagnostic tailored to this setting. Theoretical analysis and Monte Carlo simulations demonstrate that the estimator achieves higher precision under strong instruments and more robust inference under weak instruments. Empirical applications across three immigration studies confirm that the method replicates classic 2SLS findings while also detecting scenarios of weak identification, thereby supporting more cautious causal conclusions.
This study addresses the incidental parameter problem arising from unit-specific fixed effects in nonlinear panel data models, particularly when the number of time periods per unit is small and conventional estimators break down. The authors propose a novel projection-based approach that eliminates these incidental parameters without imposing assumptions on the joint distribution of fixed effects and covariates. By constructing an identified set through an implementable correspondence between observables and unobserved heterogeneity, and leveraging random set theory together with moment inequalities, they develop a distribution-free partial identification framework. This framework accommodates both static and dynamic models as well as discrete and continuous outcomes, enabling robust inference even in short panels.
This study addresses the challenge of identifying causal effects in nonseparable models when unobserved time-varying individual heterogeneity is correlated with explanatory variables. The authors propose a novel approach that approximates the conditional average potential outcomes using linear sieve methods, combined with individual-specific ridge regression and bias correction. Within a large-T asymptotic framework, this method achieves point identification, circumventing the partial identification issues common in traditional approaches. The resulting estimator admits an empirical Bayes interpretation, accommodates discrete treatment variables, and provides a unified framework for estimating average causal effects, counterfactual consumer welfare measures, and individual tax elasticities. An empirical application to supermarket scanner data demonstrates the method’s effectiveness by precisely quantifying the average equivalent variation and deadweight loss induced by price increases.