Score
Designs and analyzes algorithms, formulas, and procedures that estimate the variance or asymptotic variance of sample-based estimators and statistics, including constructing consistent and robust estimators, deriving asymptotic variance expressions, and adapting methods for quantile estimators or unbounded variables. Work covers online/updating variance estimators, proving consistency and asymptotic properties, and modifying estimators to improve finite-sample inference accuracy and robustness.
This study addresses the lack of asymptotically valid confidence intervals for quantile estimation under heterogeneous data. The authors propose a novel consistent estimator that accurately accounts for the asymptotic variance reduction induced by data grouping, thereby enabling the construction of asymptotically efficient confidence intervals. This approach provides, for the first time, theoretically guaranteed asymptotically correct confidence intervals for quantiles in heterogeneous settings, with substantially reduced interval width. Rigorous theoretical analysis establishes both the consistency of the proposed estimator and the asymptotic validity of the resulting intervals. Extensive simulations demonstrate that the constructed intervals achieve coverage probabilities close to the nominal level across various heterogeneity scenarios and are markedly shorter than those derived under the conventional i.i.d. assumption.
This paper addresses the challenge of design-based inference for the average treatment effect (ATE) in finely stratified randomized experiments—particularly under the extreme stratification regime where each stratum contains only one treated or one control unit. We propose a novel pairwise-differenced-mean variance estimator that pairs adjacent, similar strata. Unlike existing estimators, ours remains well-defined and upwardly biased with controllable magnitude even in the single-unit-per-stratum limit. Under a similarity assumption on adjacent strata, we prove analytically that our estimator exhibits reduced bias and is asymptotically superior to state-of-the-art alternatives. Finite-population bias analysis and i.i.d. superpopulation modeling, corroborated by Monte Carlo simulations, demonstrate that under high-quality stratification, our method yields substantially narrower confidence intervals and improved inferential accuracy. Our key contribution is the first variance estimation framework that simultaneously ensures theoretical rigor—via finite-sample bias characterization and asymptotic dominance—and practical robustness across realistic stratification scenarios.
This paper addresses complex randomized experiments subject to interference between units—such as social network interventions—where standard causal inference assumptions fail. Method: We develop a design-based theoretical framework for estimating treatment effects, introducing a family of design-compatible estimators and a scalar, interpretable measure of “experimental complexity.” We establish its theoretical connection to design variance, derive the asymptotic variance lower bound for unbiased estimation under arbitrary designs, and propose a consistent variance estimator. Contributions/Results: Through interference modeling, design-based inference foundations, and network experiment simulations, we validate our approach on real-world social network data from an insurance adoption study. Our estimators achieve significantly improved estimation accuracy and consistent variance estimation compared to existing methods, providing a theoretically rigorous yet practically implementable analytical framework for complex experimental designs.
This study addresses volatility estimation in high-frequency financial data contaminated by market frictions or microstructure noise. The authors propose an adaptive subsampling approach that directly infers the asymptotic (conditional) covariance matrix of volatility estimators without explicitly modeling the noise structure. By employing time-rescaled statistics over local intervals to assess sampling variability, the method automatically selects tuning parameters while ensuring the resulting covariance matrix is positive semidefinite. Theoretical analysis, Monte Carlo simulations, and empirical applications demonstrate that the proposed estimator is consistent, exhibits strong finite-sample performance, and enables robust and feasible statistical inference.
This study addresses the estimation of parameters of the form θ₀ = E[F_Y⁻¹∘F_Z(X)] in the “changes-in-changes” model, for which existing methods lack theoretical guarantees when variables are unbounded. The authors construct a plug-in estimator based on empirical quantiles and establish its √n-consistency and asymptotic normality under assumptions weaker than those in the current literature. They further propose a novel consistent estimator for the asymptotic variance. The theoretical analysis leverages empirical process theory and plug-in methods for quantile functions. Monte Carlo simulations demonstrate that the proposed variance estimator substantially outperforms existing alternatives, leading to markedly improved inference accuracy.
This study addresses the problem that standard variance estimators systematically underestimate the true variance—and consequently lead to over-rejection in hypothesis tests—when random vectors exhibit heterogeneous means and possess either clustered or weakly dependent structures. The paper proposes a novel conservative variance estimator for sums of random vectors in a triangular array setting, which is robust to mean heterogeneity and accommodates two-way clustering or weak dependence. Leveraging asymptotic theory and robust construction techniques, the proposed method maintains computational feasibility while effectively controlling test size, thereby avoiding the excessive rejection rates associated with conventional approaches and ensuring valid statistical inference.
This study addresses the inefficiency of traditional quantile estimation under light-tailed, heavy-tailed, or asymmetric distributions and its difficulty in smoothly bridging central and tail regions. The authors propose a unified interpolation-based quantile estimation framework that incorporates quadratic, Huber, or Tukey bisquare regularization into the check loss function, enabling continuous control of the effective quantile level via an interpolation parameter. They derive, for the first time, a closed-form parametrization of the effective quantile level under quadratic interpolation and establish a complete asymptotic theory, revealing the dependence of estimation efficiency on distributional shape. Theoretical and simulation results demonstrate that the proposed method reduces asymptotic variance by up to 36% under light-tailed distributions and by up to 57% under heavy-tailed or asymmetric distributions. Empirical analysis of daily log-returns confirms its superior performance in tail risk estimation.
Modern heterogeneity-robust difference-in-differences estimators derive their asymptotic properties under iid, cluster, or fixed-design frameworks that abstract from complex survey sampling, yet practitioners routinely apply them to nationally representative surveys with stratified cluster designs. We show that, under standard regularity conditions, the influence functions of each smooth IF-based or regression-based modern DiD estimator satisfy Binder's (1983) smoothness conditions, so the standard stratified-cluster variance formula applied to their values produces design-consistent standard errors. A Monte Carlo study with 66,000 replications shows where the design effect comes from. HC1 standard errors that treat observations as iid produce coverage as low as 34% under a baseline survey design and below 11% under informative sampling. Combining the survey-weighted point estimate with PSU-level clustering - the practitioner's cluster=psu heuristic - recovers near-nominal coverage across all scenarios. Adding strata and finite-population corrections yields incremental precision but is not required for valid coverage. Survey-weighted doubly robust estimation produces well-calibrated inference when parallel trends hold only conditionally. An NHANES illustration of the ACA dependent coverage provision shows that point estimates and standard errors change substantively - enough to reverse significance conclusions - when the survey design is accounted for. We provide diff-diff (https://github.com/igerber/diff-diff), an open-source Python package implementing design-based variance for fifteen modern DiD estimators.
This study addresses the substantial bias often incurred by conventional variance estimators—such as those based on Taylor linearization—for the generalized regression (GREG) estimator in high-dimensional auxiliary variable settings, which undermines inferential reliability. The paper provides the first systematic characterization of the asymptotic bias of GREG variance estimation under high-dimensional regimes and introduces a novel cross-validation–based approach to construct an unbiased variance estimator. Under mild distributional assumptions on the covariates, the proposed method is shown to be asymptotically unbiased. By integrating high-dimensional asymptotic theory with the model-assisted estimation framework, this work establishes a rigorous theoretical foundation for variance estimation in high-dimensional GREG settings and demonstrates through numerical experiments that the method exhibits excellent finite-sample performance.