🤖 AI Summary
This study addresses the lack of theoretical guarantees for $k$-nearest neighbors (kNN) regression when applied to survey data arising from complex sampling designs, which violate the standard i.i.d. assumption. The paper establishes the first consistency framework for kNN regression under such settings by integrating probability sampling theory, nonparametric regression, and asymptotic analysis. It rigorously derives a lower bound on the convergence rate and demonstrates that kNN regression remains consistent even when accounting for sampling weights and dependence structures inherent in complex surveys. However, the convergence rate is shown to be adversely affected by the curse of dimensionality. These theoretical findings extend the classical kNN consistency results—previously limited to i.i.d. data—to realistic survey contexts. The validity of the theoretical conclusions is corroborated through both simulation studies and analyses of real-world survey data.
📝 Abstract
We study the consistency of the $k$-nearest neighbor regressor under complex survey designs. While consistency results for this algorithm are well established for independent and identically distributed data, corresponding results for complex survey data are lacking. We show that the $k$-nearest neighbor regressor is consistent under regularity conditions on the sampling design and the distribution of the data. We derive lower bounds for the rate of convergence and show that these bounds exhibit the curse of dimensionality, as in the independent and identically distributed setting. Empirical studies based on simulated and real data illustrate our theoretical findings.