On Neighbourhood Cross Validation

šŸ“… 2024-04-25
šŸ“ˆ Citations: 3
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
Cross-validation (CV) for hyperparameter optimization in smoothing-penalized regression suffers from high computational cost and systematic underestimation of smoothing parameters due to unmodeled short-range autocorrelation. Method: We propose Neighborhood CV—a novel CV variant that combines QR decomposition and matrix calculus to compute CV gradients analytically and efficiently; it employs a neighborhood hold-out strategy with robust variance estimation, reducing hyperparameter optimization complexity to O(n), equivalent to a single model fit. Contribution/Results: This is the first method enabling efficient hyperparameter optimization and uncertainty quantification via neighborhood CV under short-range autocorrelation. Experiments demonstrate substantial improvements in accuracy and robustness of smoothing parameter estimation across smoothing quantile regression and GAMLSS models. Neighborhood CV establishes a scalable, theoretically grounded, and practical paradigm for hyperparameter selection in high-dimensional nonparametric modeling.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Learning & Optimization for CVSearch and Optimization: Non-convex Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
šŸ“ Abstract
Many varieties of cross validation would be statistically appealing for the estimation of smoothing and other penalized regression hyperparameters, were it not for the high cost of evaluating such criteria. Here it is shown how to efficiently and accurately compute and optimize a broad variety of cross validation criteria for a wide range of models estimated by minimizing a quadratically penalized loss. The leading order computational cost of hyperparameter estimation is made comparable to the cost of a single model fit given hyperparameters. In many cases this represents an $O(n)$ computational saving when modelling $n$ data. This development makes if feasible, for the first time, to use leave-out-neighbourhood cross validation to deal with the wide spread problem of un-modelled short range autocorrelation which otherwise leads to underestimation of smoothing parameters. It is also shown how to accurately quantifying uncertainty in this case, despite the un-modelled autocorrelation. Practical examples are provided including smooth quantile regression, generalized additive models for location scale and shape, and focussing particularly on dealing with un-modelled autocorrelation.
Problem

Research questions and friction points this paper is trying to address.

Efficiently compute cross validation for penalized regression hyperparameters
Reduce computational cost to that of a single model fit
Address un-modelled short range autocorrelation in smoothing parameters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Efficient computation of cross validation criteria
Optimization with cost comparable to single model fit
Handling un-modelled autocorrelation via neighbourhood validation
šŸ”Ž Similar Papers
No similar papers found.
šŸ’¼ Related Jobs
No related jobs found.
University of Edinburgh
S
Simon N. Wood
School of Mathematics, University of Edinburgh, U.K.