Score
Designs and implements two-stage estimation procedures that propagate first-stage model parameters (for example slopes or random-effect estimates) into a second-stage model while applying a formal correction step to remove the bias introduced by naive plug‑in approaches. Builds algorithms and variance-propagation methods that re‑estimate or adjust slope parameters and other quantities to reduce bias compared with standard two‑stage estimators while keeping the computational efficiency of two‑stage workflows.
该研究通过将训练子样本视为第二阶段抽样,解决了模型辅助估计中的不确定性量化问题,并提出了两种方差估计方法。
Two-stage factor score regression (FSR) often yields biased structural parameter estimates due to the use of estimated factor scores. This work proposes a general bias-correction framework applicable to a broad class of parametric latent variable models, obviating the need to explicitly compute model-specific factor scores. The approach overcomes limitations of existing correction techniques and achieves √n-consistency under mild regularity conditions. Point estimates are obtained via a stochastic approximation algorithm, while variance estimation leverages Monte Carlo simulation, substantially reducing reliance on intricate analytical derivations. Simulation studies demonstrate that the proposed method attains estimation accuracy comparable to one-stage maximum likelihood estimation—the “gold standard”—offering a simple yet efficient alternative for two-stage modeling.
Under network interference, causal effect estimation suffers from high variance, while simultaneously minimizing cut edges within clusters and achieving covariate balance remains challenging. Method: We propose a two-stage rollout experimental design: (1) graph-based clustering to identify highly homogeneous subpopulations, followed by (2) intervention deployment exclusively within those subpopulations. Contribution/Results: We formally link clustering objectives—cut-edge minimization versus covariate balance—to the bias–variance trade-off in causal estimation, theoretically characterizing how cluster structure affects bias (governed by cut edges) and variance (driven by homogeneity and covariate balance). Using a polynomial interpolation estimator and Monte Carlo simulations, we empirically identify optimal trade-offs across diverse clustering strategies. Our approach significantly reduces estimation variance while preserving causal identification validity under interference.
This paper addresses the challenge of design-based inference for the average treatment effect (ATE) in finely stratified randomized experiments—particularly under the extreme stratification regime where each stratum contains only one treated or one control unit. We propose a novel pairwise-differenced-mean variance estimator that pairs adjacent, similar strata. Unlike existing estimators, ours remains well-defined and upwardly biased with controllable magnitude even in the single-unit-per-stratum limit. Under a similarity assumption on adjacent strata, we prove analytically that our estimator exhibits reduced bias and is asymptotically superior to state-of-the-art alternatives. Finite-population bias analysis and i.i.d. superpopulation modeling, corroborated by Monte Carlo simulations, demonstrate that under high-quality stratification, our method yields substantially narrower confidence intervals and improved inferential accuracy. Our key contribution is the first variance estimation framework that simultaneously ensures theoretical rigor—via finite-sample bias characterization and asymptotic dominance—and practical robustness across realistic stratification scenarios.
Accurate estimation and inference for the average treatment effect (ATE) in multi-covariate randomized controlled trials (RCTs) remain challenging, particularly under high-dimensional covariates and small sample sizes. Method: Building on the Neyman finite-population framework, we propose a bias-corrected regression-adjustment estimator with cross-fitting and introduce, for the first time, an HC3-type heteroskedasticity-robust standard error tailored to high-dimensional settings. We rigorously derive first- and second-order stochastic expansions of the random component of regression-adjustment estimators, identifying the source of higher-order bias in conventional inference; leveraging this insight, we design a cross-fitting procedure to eliminate bias and extend HC3 standard errors to stratified experimental designs. Results: Simulations and reanalysis of Angrist et al. (2009)’s education RCT demonstrate that our method substantially improves estimation accuracy, confidence interval coverage, and statistical power—especially in small-sample and high-dimensional scenarios.
This study addresses efficiency losses and inferential challenges in two-sample instrumental variable estimation arising from sample heterogeneity, heteroskedasticity, weak instruments, and overidentification. The authors propose a robust two-step estimator that relies solely on summary statistics—namely, coefficient vectors and variance matrices—from the reduced-form and first-stage regressions in each sample. By relaxing conventional homoskedasticity and homogeneity assumptions, the method achieves efficient estimation under heteroskedasticity and heterogeneity. It extends the Montiel-Olea-Pflueger effective F-statistic to diagnose weak instruments in the two-sample context and develops a Hansen-type overidentification test tailored to this setting. An empirical application examining the causal effect of education on voting behavior demonstrates the method’s validity and practical utility, enabling reliable causal inference with complex data structures.
本文提出了一种学习者无关的框架,通过结合机器学习和设计意识交叉拟合来改进调查数据中有限总体参数估计的有效性。
This study addresses the challenge of data missingness in longitudinal two-stage studies due to participant dropout and outcome subsampling, where conventional inverse probability weighting methods suffer from low efficiency and fail to leverage covariate information. The authors propose two novel approaches by integrating such designs into the longitudinal targeted maximum likelihood estimation (LTMLE) framework: first, an IPCW-LTMLE that incorporates known sampling weights, and second, an unweighted LTMLE that treats the sampling indicator as an intervention node within sequential regression. Theoretical and simulation results demonstrate that the proposed LTMLE methods reduce variance by up to 73% compared to weighted Kaplan–Meier estimators—typically achieving 30–50% gains—with IPCW-LTMLE further improving efficiency by 20–35%. When combined with cross-fitted variance estimation, nominal confidence interval coverage is restored from as low as 76% to the desired level, substantially enhancing inferential validity.
This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.
This study addresses unbiased and robust variance estimation for average treatment effects in finely stratified experiments where only one unit per stratum receives treatment. The authors propose a graph Laplacian–based variance estimator that treats strata as vertices in a graph and aggregates cross-stratum information through edge weights, thereby unifying classical approaches such as paired and complete-graph estimators. Theoretical analysis reveals that bias arises from differences in treatment effects between adjacent strata, and establishes an exact bias identity for degree-calibrated estimators. Building on this insight, they design a regularized graph estimator that balances locality with worst-case robustness. The complete-graph estimator is shown to achieve minimax normalized bias under weak heterogeneity, while the regularized variant effectively trades off bias control and local adaptivity in simulations.