Genuinely Robust Inference for Clustered Data

📅 2023-08-20
📈 Citations: 2
Influential: 0
📄 PDF

career value

216K/year
🤖 AI Summary
Conventional clustered robust inference fails when cluster sizes are non-negligible—e.g., following Zipf’s law—and 77% of empirical studies in the *American Economic Review* and *Econometrica* (2020–2021) violate its implicit equal-size or bounded-size assumptions. Method: This paper establishes the first necessary and sufficient condition for consistency of clustered robust estimators and proposes two new procedures: score subsampling and size-adjusted reweighting. Both methods are theoretically grounded—guaranteeing consistency and uniform size control—and practically implementable, with ready-to-use Stata packages. Results: Monte Carlo simulations demonstrate that the proposed methods strictly maintain nominal test size even where conventional approaches severely distort inference. They constitute the first truly robust and implementable inferential framework for settings with large, heterogeneous cluster sizes.
📝 Abstract
Conventional cluster-robust inference methods are inconsistent when clusters are unignorably large. We derive a necessary and sufficient condition for consistency, which is violated in 77% of empirical studies published in American Economic Review and Econometrica (2020-2021). To address this, we propose two methods: (i) score subsampling, which retains the original estimator, and (ii) size-adjusted reweighting, which is easy to implement in software like Stata and remains valid if the cluster size follows Zipf's law. Simulations confirm the reliability and uniform size control of these approaches, offering robust alternatives where conventional methods fail.
Problem

Research questions and friction points this paper is trying to address.

Conventional cluster-robust inference fails with large clusters
A new condition reveals frequent inconsistency in published research
Proposes a novel bootstrap method for valid inference across processes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel cluster score bootstrap method
Valid inference for large cluster sizes
Robust size control across data processes
🔎 Similar Papers