🤖 AI Summary
This study addresses the limitation of traditional ridge regression, which employs a single global regularization parameter and thus struggles to achieve individualized predictions. We propose Focused Ridge Regression (Fridge), a method that estimates a unique optimal tuning parameter for each covariate vector. Specifically, we define an oracle parameter that minimizes the individual mean squared error and develop a plug-in estimator to approximate it. The framework is further extended to logistic regression and high-dimensional settings, with an accompanying R package, fridge, implemented for practical use. Simulation studies and real-world medical data analyses demonstrate that Fridge yields significantly lower average prediction errors compared to conventional cross-validated ridge regression, effectively enhancing the accuracy of personalized risk prediction.
📝 Abstract
Statistical prediction methods typically require some form of fine-tuning of tuning parameter(s), with $K$-fold cross-validation as the canonical procedure. For ridge regression there exist numerous procedures, but common for all, including cross-validation, is that one single parameter is chosen for all future predictions. We propose instead to calculate a unique tuning parameter for each individual for which we wish to predict an outcome. This generates an individualized prediction by focusing on the vector of covariates of a specific individual. The focused ridge -- fridge -- procedure is introduced with a two-part contribution: 1) first we define an oracle tuning parameter minimizing the mean squared prediction error of a specific covariate vector, 2) then we propose to estimate this tuning parameter by using plug-in estimates of the regression coefficients and error variance parameter. The procedure is extended to logistic ridge regression by utilizing parametric bootstrap. For high-dimensional data, we propose to use ridge regression with cross-validation as the plug-in estimate, and simulations show that fridge gives smaller average prediction error than ridge with cross-validation for both simulated and real data. We illustrate the new concept for both linear and logistic regression models in two applications of personalized medicine: predicting individual risk and treatment response based on gene expression data. The method is implemented in the R package "fridge".