🤖 AI Summary
In high-dimensional network analysis, when the number of observations is much smaller than the number of variables, model misspecification often leads to an inflated false positive rate in neighborhood estimation. This work proposes a conservative neighborhood selection method based on penalizing the volume of the model space, extending the Minimum Description Length (MDL) principle to high-dimensional settings and integrating ridge regression to elucidate its impact on mean squared error. The approach is applicable under both linear and nonlinear true models and theoretically guarantees either consistent neighborhood recovery under correct specification or a sparser—yet safer—estimate under misspecification, thereby substantially reducing false positive edges. This strategy overcomes key limitations of conventional methods such as Lasso and information criteria like AIC and BIC.
📝 Abstract
To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a conservative (i.e., low false positive rate) estimate of the neighbourhood for each node (which connections are in the network). A neighbourhood is often estimated with a linear model, and this leads to two interesting cases: (i) If the true model is linear, then neighbourhood selection work reasonably well, and (ii) if the true model is nonlinear, then neighbourhood selection requires a penalty for the high dimensions. Here we show the impact of the ridge parameter on the mean squared error, and how this leads to low test variance and hence to neighbourhoods with large numbers of edges. We connect these insights with results from machine learning, where the so-called double descent (when more parameters are included than observations, the mean squared error goes down a second time) has put the traditional view on model selection upside down. Essentially, for adequate neighbourhood selection in models with a large number of parameters, the volume of the model space needs to be included in the penalty. Most neighbourhood selection methods (e.g., Lasso, AIC, BIC) lead to spurious edges (high false positive rate), but we prove that in the high-dimensional setting, minimum description length leads to correct neighbourhood selection or smaller (low false positive rates) in both cases when either the model is correctly or incorrectly assumed linear