Score
Designs and evaluates statistical estimators and procedures for measuring the proportion or frequency of a characteristic in a population from observed data, including sampling schemes, weighting, and adjustments for missingness or measurement error. Builds estimation algorithms, computes uncertainty (confidence intervals, variance estimates), and analyzes estimator properties such as bias, variance, consistency and robustness under complex sampling and imperfect measurement.
This study addresses the limitations of conventional inference methods rooted in sampling variability when sample sizes approach the population size. By constructing finite populations with known parameters and leveraging CPU/GPU-accelerated repeated sampling experiments, the authors examine the evolution of the randomization distribution of the sample mean across varying sampling fractions. Integrating finite population theory with numerical precision analysis, they demonstrate that in high-coverage scenarios, estimation error predominantly stems from computational precision and architectural constraints rather than sampling randomness. The findings reveal that sampling variability becomes negligible well before exhaustive enumeration is reached, thereby challenging a foundational assumption of classical inferential statistics and offering a basis for rethinking statistical paradigms in the context of large-scale, near-complete data.
This study addresses the inconsistency of conventional estimators in mixed-data sampling (MIDAS) regression when both high- and low-frequency variables are subject to measurement error. To resolve this issue, the paper introduces the corrected score method into the MIDAS framework for the first time and combines it with profile likelihood to construct a consistent estimator. This approach effectively overcomes the inconsistency that plagues existing profile likelihood estimators under measurement error. Through comprehensive Monte Carlo simulations, the authors systematically investigate the impacts of sample size, lag order, and nuisance parameters on estimation performance. The results demonstrate that the proposed estimator exhibits strong consistency and favorable finite-sample properties across a range of sample sizes and model specifications.
To address nonignorable nonresponse in sample surveys, this paper extends model-assisted estimation to the missing-at-random (MAR) framework. We propose a calibratable inverse-probability weighting (IPW) method that reweights sampled units in a second stage to compensate for nonrespondents, and systematically construct a Horvitz–Thompson-type adjusted estimator. Theoretically, we establish its asymptotic design-unbiasedness and design-consistency, derive a closed-form asymptotic variance expression, and provide a consistent variance estimator. Monte Carlo simulations demonstrate that the proposed estimator significantly outperforms the conventional Horvitz–Thompson estimator under diverse nonresponse mechanisms. Our key contributions are: (i) the first systematic adaptation of model-assisted estimation to the MAR setting; and (ii) a novel IPW weighting scheme that simultaneously satisfies calibration constraints and enjoys rigorous asymptotic properties—namely, design-consistency, asymptotic normality, and consistent variance estimation.
This paper addresses the challenge of design-based inference for the average treatment effect (ATE) in finely stratified randomized experiments—particularly under the extreme stratification regime where each stratum contains only one treated or one control unit. We propose a novel pairwise-differenced-mean variance estimator that pairs adjacent, similar strata. Unlike existing estimators, ours remains well-defined and upwardly biased with controllable magnitude even in the single-unit-per-stratum limit. Under a similarity assumption on adjacent strata, we prove analytically that our estimator exhibits reduced bias and is asymptotically superior to state-of-the-art alternatives. Finite-population bias analysis and i.i.d. superpopulation modeling, corroborated by Monte Carlo simulations, demonstrate that under high-quality stratification, our method yields substantially narrower confidence intervals and improved inferential accuracy. Our key contribution is the first variance estimation framework that simultaneously ensures theoretical rigor—via finite-sample bias characterization and asymptotic dominance—and practical robustness across realistic stratification scenarios.
This study addresses the coarsening of self-reported numeric variables in surveys—often caused by rounding or heaping—by proposing a novel approach that integrates design-based inference with latent variable modeling. Treating observed values as coarsened manifestations of an underlying continuous latent variable, the method jointly models the coarsening mechanism and the latent distribution via a survey-weighted pseudo-likelihood. It generates posterior predictive replicates to propagate coarsening-induced uncertainty into standard design-based estimators. This framework is the first to explicitly correct for coarsening bias under complex sampling designs, enabling unbiased estimation of means, quantiles, and threshold-based prevalence measures. Simulation studies demonstrate robustness across various model misspecifications and sampling scenarios, and empirical application to Italy’s PASSI behavioral surveillance data shows effective correction of coarsening-related estimation bias.
This study addresses the lack of effective methods for quantifying how observables in singular statistical models respond to data perturbations. It introduces susceptibility as a central measure of this response and proposes, for the first time, an estimator for generalized observables grounded in linear response theory. By integrating statistical inference with asymptotic analysis, the estimator is shown to be consistent and asymptotically unbiased in the large-sample limit, relying solely on $n$ observed data points. This work establishes the first theoretically guaranteed framework for sensitivity analysis in singular statistical models, providing rigorous foundations for assessing the stability of model outputs under infinitesimal data perturbations.
Modern heterogeneity-robust difference-in-differences estimators derive their asymptotic properties under iid, cluster, or fixed-design frameworks that abstract from complex survey sampling, yet practitioners routinely apply them to nationally representative surveys with stratified cluster designs. We show that, under standard regularity conditions, the influence functions of each smooth IF-based or regression-based modern DiD estimator satisfy Binder's (1983) smoothness conditions, so the standard stratified-cluster variance formula applied to their values produces design-consistent standard errors. A Monte Carlo study with 66,000 replications shows where the design effect comes from. HC1 standard errors that treat observations as iid produce coverage as low as 34% under a baseline survey design and below 11% under informative sampling. Combining the survey-weighted point estimate with PSU-level clustering - the practitioner's cluster=psu heuristic - recovers near-nominal coverage across all scenarios. Adding strata and finite-population corrections yields incremental precision but is not required for valid coverage. Survey-weighted doubly robust estimation produces well-calibrated inference when parallel trends hold only conditionally. An NHANES illustration of the ACA dependent coverage provision shows that point estimates and standard errors change substantively - enough to reverse significance conclusions - when the survey design is accounted for. We provide diff-diff (https://github.com/igerber/diff-diff), an open-source Python package implementing design-based variance for fifteen modern DiD estimators.
This study addresses the issue of variance inflation in regression models under complex survey designs, which often arises from unnecessary variability in sampling weights. The authors propose a novel approach that, for the first time, integrates stabilized weights with generalized raking within a two-stage sampling framework, leveraging auxiliary covariate information to effectively reduce extraneous weight variation. This method substantially enhances the efficiency of design-based estimators while remaining compatible with standard statistical software. Simulation studies demonstrate that, under typical two-stage survey designs, the proposed estimator achieves markedly higher precision compared to existing methods. The approach has been successfully applied to a large-scale multinational study of Kaposi’s sarcoma, illustrating its practical utility and robustness in real-world settings.