Bandwidth Selection for Spatial HAC Standard Errors

📅 2026-03-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underestimation of regression standard errors caused by spatial autocorrelation and the lack of systematic guidance for bandwidth selection in existing spatial HAC (heteroskedasticity and autocorrelation consistent) methods. The authors propose a nonparametric, data-driven bandwidth selection procedure based on the empirical covariogram of residuals for spatial HAC standard error estimation. Their analysis reveals an inverted U-shaped relationship between bandwidth and standard errors, challenging the conventional wisdom that larger bandwidths are inherently more conservative, and provides the first formalized bandwidth selection framework. Monte Carlo simulations across diverse spatial structures and sample configurations demonstrate that the method—particularly with Bartlett or Epanechnikov kernels—maintains empirical size close to the nominal 5% level. The approach is further validated using U.S. county-level data, and the accompanying R package SpatialInference is publicly available.

Technology Category

Machine Learning: Kernel MethodsData Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal DataReasoning under Uncertainty: Stochastic Optimization

Application Category

Security and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Spatial autocorrelation in regression models can lead to downward biased standard errors and thus incorrect inference. The most common correction in applied economics is the spatial heteroskedasticity and autocorrelation consistent (HAC) standard error estimator introduced by Conley (1999). A critical input is the kernel bandwidth: the distance within which residuals are allowed to be correlated. However, this is still an unresolved problem and there is no formal guidance in the literature. In this paper, I first document that the relationship between the bandwidth and the magnitude of spatial HAC standard errors is inverse-U shaped. This implies that both too narrow and too wide bandwidths lead to underestimated standard errors, contradicting the conventional wisdom that wider bandwidths yield more conservative inference. I then propose a simple, non-parametric, data-driven bandwidth selector based on the empirical covariogram of regression residuals. In extensive Monte Carlo experiments calibrated to empirically relevant spatial correlation structures across the contiguous United States, I show that the proposed method controls the false positive rate at or near the nominal 5% level across a wide range of spatial correlation intensities and sample configurations. I compare six kernel functions and find that the Bartlett and Epanechnikov kernels deliver the best size control. An empirical application using U.S. county-level data illustrates the practical relevance of the method. The R package SpatialInference implements the proposed bandwidth selection method.
Problem

Research questions and friction points this paper is trying to address.

bandwidth selection
spatial HAC standard errors
spatial autocorrelation
kernel bandwidth
standard error estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatial HAC
bandwidth selection
empirical covariogram
data-driven
Monte Carlo simulation
💼 Related Jobs
No related jobs found.