🤖 AI Summary
Conventional filter-based variable selection relies solely on marginal correlations, neglecting dependencies among predictors and thus failing to identify synergistic or context-dependent effects. Method: This paper pioneers the integration of relative importance (RI) analysis as a preprocessing step for variable ranking and screening. We propose CRI.Z—a computationally efficient method that combines generalized dominance (GD) analysis with composite relative importance (CRI) to decompose and quantify both direct and joint contributions of predictors within a multiple regression framework. Contribution/Results: CRI.Z effectively identifies key variables within highly correlated clusters and detects weak-margin but high-synergy predictors. Empirical evaluation demonstrates that RI-based filtering substantially outperforms Lasso and Relaxed Lasso in high-dimensional, multicollinear settings—yielding improved prediction accuracy and model stability. The approach establishes a novel, interpretable paradigm for variable selection grounded in effect decomposition rather than sparsity alone.
📝 Abstract
Although conceptually related, variable selection and relative importance (RI) analysis have been treated quite differently in the literature. While RI is typically used for post-hoc model explanation, this paper explores its potential for variable ranking and filter-based selection before model creation. Specifically, we anticipate strong performance from the RI measures because they incorporate both direct and combined effects of predictors, addressing a key limitation of marginal correlation that ignores dependencies among predictors. We implement and evaluate the RI-based variable selection methods using general dominance (GD), comprehensive relative importance (CRI), and a newly proposed, computationally efficient variant termed CRI.Z.
We first demonstrate how the RI measures more accurately rank the variables than the marginal correlation, especially when there are suppressed or weak predictors. We then show that predictive models built on these rankings are highly competitive, often outperforming state-of-the-art methods such as the lasso and relaxed lasso. The proposed RI-based methods are particularly effective in challenging cases involving clusters of highly correlated predictors, a setting known to cause failures in many benchmark methods. Although lasso methods have dominated the recent literature on variable selection, our study reveals that the RI-based method is a powerful and competitive alternative. We believe these underutilized tools deserve greater attention in statistics and machine learning communities. The code is available at: https://github.com/tien-endotchang/RI-variable-selection.