🤖 AI Summary
This work addresses the optimal train-test split for ridge regression in the high-dimensional asymptotic regime where both sample size $m$ and feature dimension $n$ diverge, with $n/m o gamma in (0,1)$. The objective is to maximize “completeness” of model evaluation—i.e., minimize the asymptotic bias between test error and theoretical generalization error. We propose the first rigorous large-sample analytical framework for determining the optimal split ratio, leveraging high-dimensional asymptotics and random matrix theory to optimize the ridge regression error functional. Our analysis reveals that the optimal split depends asymptotically only on $m$ and $n$, and is nearly insensitive to the regularization parameter $alpha$. Moreover, its first two asymptotic expansion terms coincide with those of ordinary linear regression, rendering it practically parameter-free. This yields the first theoretically grounded, optimal data allocation principle for model evaluation in high dimensions.
📝 Abstract
We derive the ideal train/test split for the ridge regression to high accuracy in the limit that the number of training rows m becomes large. The split must depend on the ridge tuning parameter, alpha, but we find that the dependence is weak and can asymptotically be ignored; all parameters vanish except for m and the number of features, n. This is the first time that such a split is calculated mathematically for a machine learning model in the large data limit. The goal of the calculations is to maximize"integrity,"so that the measured error in the trained model is as close as possible to what it theoretically should be. This paper's result for the ridge regression split matches prior art for the plain vanilla linear regression split to the first two terms asymptotically, and it appears that practically there is no difference.