event-counting estimation

Designs and analyzes estimators that use counts of observed indicator events under full observation to estimate statistical quantities (e.g., mixture proportions). This work builds counting-based or full-observation estimators and characterizes their bias, variance and sample complexity, often including proofs of matching information-theoretic lower bounds.

event-countingestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.71
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Design-based Estimation Theory for Complex Experiments

Nov 12, 2023
HC
Haoge Chang
🏛️ Columbia University

This paper addresses complex randomized experiments subject to interference between units—such as social network interventions—where standard causal inference assumptions fail. Method: We develop a design-based theoretical framework for estimating treatment effects, introducing a family of design-compatible estimators and a scalar, interpretable measure of “experimental complexity.” We establish its theoretical connection to design variance, derive the asymptotic variance lower bound for unbiased estimation under arbitrary designs, and propose a consistent variance estimator. Contributions/Results: Through interference modeling, design-based inference foundations, and network experiment simulations, we validate our approach on real-world social network data from an insurance adoption study. Our estimators achieve significantly improved estimation accuracy and consistent variance estimation compared to existing methods, providing a theoretically rigorous yet practically implementable analytical framework for complex experimental designs.

Developing design-based estimation theory for arbitrary designsEstimating treatment effects in complex randomized experimentsProposing new estimators with favorable asymptotic properties

This paper addresses systematic bias in Theil, Atkinson, and discrete inequality index estimators when the underlying population follows a finite gamma mixture distribution. We propose an analytical framework to derive closed-form bias expressions for these three inequality measures under heterogeneous gamma mixture models. Leveraging Mosimann’s proportionality and independence theorem, together with the intrinsic relationship between gamma and Dirichlet distributions, we obtain exact bias formulas—marking the first such derivation for mixed gamma settings. Unlike prior work restricted to single gamma assumptions, our approach enables precise quantification of estimation bias in nonhomogeneous populations. The resulting explicit analytical expressions substantially improve the statistical accuracy and theoretical applicability of inequality measurement in empirically heterogeneous contexts—such as income or health distributions—thereby providing a more robust econometric foundation for inequality analysis.

Analyzing gamma mixture populations with closed-form expressionsEstimating bias in Theil, Atkinson, and dispersion indicesExtending single gamma model results to heterogeneous populations

A New Design-Based Variance Estimator for Finely Stratified Experiments

Mar 13, 2025
YB
Yuehao Bai
🏛️ University of Southern California | University of Chicago | Stanford University

This paper addresses the challenge of design-based inference for the average treatment effect (ATE) in finely stratified randomized experiments—particularly under the extreme stratification regime where each stratum contains only one treated or one control unit. We propose a novel pairwise-differenced-mean variance estimator that pairs adjacent, similar strata. Unlike existing estimators, ours remains well-defined and upwardly biased with controllable magnitude even in the single-unit-per-stratum limit. Under a similarity assumption on adjacent strata, we prove analytically that our estimator exhibits reduced bias and is asymptotically superior to state-of-the-art alternatives. Finite-population bias analysis and i.i.d. superpopulation modeling, corroborated by Monte Carlo simulations, demonstrate that under high-quality stratification, our method yields substantially narrower confidence intervals and improved inferential accuracy. Our key contribution is the first variance estimation framework that simultaneously ensures theoretical rigor—via finite-sample bias characterization and asymptotic dominance—and practical robustness across realistic stratification scenarios.

Comparing bias with existing variance estimatorsEstimating variance in finely stratified experimentsImproving inference precision with novel estimator

Exact Sampling of Gibbs Measures with Estimated Losses

Apr 24, 2024
DF
David Frazier
🏛️ Monash University | University College London | Queensland University of Technology

This work addresses the slow MCMC convergence in Gibbs posterior sampling under stochastic loss functions, which stems from spurious dependence on the number of pseudo-observations. We propose the first pseudo-sample-size–independent corrected piecewise deterministic Markov process (PDMP) sampler. By designing a novel jump-rate function and direction mechanism, our method rigorously ensures that the invariant measure remains invariant to the pseudo-observation count—thereby overcoming the inherent trade-off between asymptotic bias and slow convergence in conventional stochastic-loss inference. We prove that the sampler converges exactly to the target Gibbs posterior measure with a uniform convergence rate independent of pseudo-sample size. Empirical validation across three canonical settings—likelihood-intractable models, misspecified models, and stochastic losses—demonstrates elimination of pseudo-sample-size bias in posterior sampling, alongside substantial improvements in robustness and estimation accuracy.

Addressing slow convergence in Gibbs measures with estimated lossesImproving inference for intractable likelihoods and model misspecificationReducing pseudo-observation dependence in MCMC posterior sampling

Finite Population Survey Sampling: An Unapologetic Bayesian Perspective.

Jun 18, 2023
SB
Sudipto Banerjee
🏛️ UCLA | University of California Los Angeles

Bayesian inference for finite-population surveys is challenging when sampling units exhibit complex dependencies (e.g., spatial, network, or structural) and nonresponse is nonignorable. Method: We propose a unified hierarchical modeling framework that integrates graphical models and spatial random fields to characterize multivariate dependence; formally adopts the “unapologetic Bayesian” paradigm, embedding design-based weights (e.g., Horvitz–Thompson) naturally into prior and likelihood specifications; incorporates causal ignorability analysis to ensure identifiability under missing-not-at-random (MNAR) mechanisms; and employs MCMC and variational inference for scalable computation. Contribution/Results: The framework achieves improved small-area estimation accuracy and more reliable uncertainty quantification in two empirical spatial finite-population analyses. It rigorously reconciles design-based consistency with model-based flexibility, providing theoretical guarantees for valid Bayesian inference under complex survey designs and nonignorable nonresponse.

Developing Bayesian frameworks for ignorable and nonignorable responsesIncorporating multivariate dependencies using graphical and spatial modelsModeling complex dependencies in finite population sampling

Latest Papers

What's happening recently
View more

This study addresses the joint optimization of sampling design and estimation under bounded population values, aiming to achieve design-unbiased estimation of the population total while minimizing the worst-case mean squared error. Within the Horvitz–Thompson framework and adopting a minimax criterion, the paper establishes—for the first time—the minimax lower bound over all design-unbiased estimators and shows that this bound is attainable when the unit inclusion indicators are pairwise independent. The authors further propose a midpoint-differenced Horvitz–Thompson estimator, which achieves minimax optimality under an independent sampling strategy with inclusion probabilities πᵢ* = min(1, c(bᵢ − aᵢ)). This estimator is also shown to be admissible within the class of unbiased and affine-equivariant estimators, thereby extending Gabler’s (1990) linear result to a broader class of estimators.

bounded outcomesdesign-unbiased estimationfinite population

This study addresses the challenge of coarsened data arising in two-stage sampling when only a subset of variables is observed in the second stage. Under the assumption that the outcome variable is fully observed, the authors propose a class of novel estimators based on targeted maximum likelihood estimation (TMLE). This approach provides a unified framework for modeling the second-stage sampling mechanism, encompassing generalized calibration estimation, inverse probability of censoring weighted TMLE (IPCW-TMLE), and their extensions. The proposed estimators possess double robustness and achieve higher efficiency, with theoretical analysis demonstrating that they attain the semiparametric efficiency bound asymptotically—matching the best-known performance in the literature—and thereby substantially improving the precision of parameter estimation.

coarsened dataefficient estimationsampling mechanism

This study addresses the coarsening of self-reported numeric variables in surveys—often caused by rounding or heaping—by proposing a novel approach that integrates design-based inference with latent variable modeling. Treating observed values as coarsened manifestations of an underlying continuous latent variable, the method jointly models the coarsening mechanism and the latent distribution via a survey-weighted pseudo-likelihood. It generates posterior predictive replicates to propagate coarsening-induced uncertainty into standard design-based estimators. This framework is the first to explicitly correct for coarsening bias under complex sampling designs, enabling unbiased estimation of means, quantiles, and threshold-based prevalence measures. Simulation studies demonstrate robustness across various model misspecifications and sampling scenarios, and empirical application to Italy’s PASSI behavioral surveillance data shows effective correction of coarsening-related estimation bias.

coarseningdesign-based estimationfinite-population inference

This study addresses the inadequate estimation accuracy of the relative risk (RR), odds ratio (OR), and their logarithmic forms for rare binary attributes in two populations. To overcome this limitation, the authors propose a sequential equal-allocation sampling method that efficiently estimates these parameters while ensuring the relative mean squared error (for RR/OR) or mean squared error (for the log-transformed parameters) remains below a pre-specified threshold. Under rare or moderately rare event settings, the proposed estimator achieves performance approaching the Cramér–Rao lower bound, offering both high efficiency and rigorous error control. The key innovation lies in integrating a sequential sampling strategy with explicit mean squared error constraints, substantially enhancing the precision and reliability of RR and OR estimation in scenarios involving rare events.

mean-square errorodds ratiorare events

Hot Scholars

CB

Chenjia Bai

Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院, TeleAI)
Reinforcement LearningRoboticsEmbodied AI
NJ

Nils Jansen

Professor of Artificial Intelligence and Formal Methods, Ruhr-University Bochum
AIDecision-Making under UncertaintyPOMDPsSafe Reinforcement Learning
SY

Siyuan Yang

Wallenberg-NTU Presidential Postdoctoral Fellowship, Nanyang Technological University
Computer VisionAction Recognition
PF

Pascal Fua

Professor Computer Science, EPFL
Computer VisionMachine LearningComputer Asisted Eng.Biomedical Imaging
AB

Antoni B. Chan

Professor of Computer Science, City University of Hong Kong
Computer VisionMachine LearningSurveillanceEye Gaze Analysis