🤖 AI Summary
This paper addresses the challenge of statistical inference for kernel ridge regression (KRR) on non-standard data—such as preference rankings, graphs, and sequences—where conventional inferential frameworks fail.
Method: We construct the first uniform confidence set for KRR with nearly minimax-optimal shrinkage rate. Our approach leverages a symmetric bootstrap procedure that automatically cancels bias while ensuring computational efficiency. Rigorous finite-sample coverage is established via reproducing kernel Hilbert space (RKHS) theory, uniform Gaussian/bootstrap coupling, and covering number analysis.
Contribution/Results: The method enables direct hypothesis testing of matching effects—for instance, whether students benefit from attending higher-ranked schools—and is empirically validated in evaluating school assignment mechanisms. By providing valid, distribution-free inference for KRR on structured domains, our work fills a critical theoretical gap in econometrics and related fields, where KRR has seen widespread application but lacked formal inferential foundations.
📝 Abstract
We provide uniform inference and confidence bands for kernel ridge regression (KRR), a widely-used non-parametric regression estimator for general data types including rankings, images, and graphs. Despite the prevalence of these data -- e.g., ranked preference lists in school assignment -- the inferential theory of KRR is not fully known, limiting its role in economics and other scientific domains. We construct sharp, uniform confidence sets for KRR, which shrink at nearly the minimax rate, for general regressors. To conduct inference, we develop an efficient bootstrap procedure that uses symmetrization to cancel bias and limit computational overhead. To justify the procedure, we derive finite-sample, uniform Gaussian and bootstrap couplings for partial sums in a reproducing kernel Hilbert space (RKHS). These imply strong approximation for empirical processes indexed by the RKHS unit ball with logarithmic dependence on the covering number. Simulations verify coverage. We use our procedure to construct a novel test for match effects in school assignment, an important question in education economics with consequences for school choice reforms.