Solving the Best Subset Selection Problem via Suboptimal Algorithms

📅 2025-03-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The NP-hard problem of best-subset selection in high-dimensional linear regression motivates this work. We propose an efficient, suboptimal algorithmic framework that integrates greedy search, regularized path tracking, and cross-validation-based model evaluation to yield a stable and scalable solution pipeline. Compared with mainstream heuristic approaches—including LASSO and orthogonal matching pursuit (OMP)—our method substantially reduces computational cost in ultra-high-dimensional settings (p ≥ 1000) while achieving superior trade-offs between model sparsity and predictive accuracy. Comprehensive benchmark experiments on both synthetic data and diverse real-world datasets demonstrate that our approach consistently attains higher solution quality and greater robustness than state-of-the-art baselines. By bridging efficiency and statistical reliability, the proposed framework establishes a new paradigm for high-dimensional sparse modeling.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionSearch and Optimization: Non-convex OptimizationReasoning under Uncertainty: Stochastic Optimization

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
Best subset selection in linear regression is well known to be nonconvex and computationally challenging to solve, as the number of possible subsets grows rapidly with increasing dimensionality of the problem. As a result, finding the global optimal solution via an exact optimization method for a problem with dimensions of 1000s may take an impractical amount of CPU time. This suggests the importance of finding suboptimal procedures that can provide good approximate solutions using much less computational effort than exact methods. In this work, we introduce a new procedure and compare it with other popular suboptimal algorithms to solve the best subset selection problem. Extensive computational experiments using synthetic and real data have been performed. The results provide insights into the performance of these methods in different data settings. The new procedure is observed to be a competitive suboptimal algorithm for solving the best subset selection problem for high-dimensional data.
Problem

Research questions and friction points this paper is trying to address.

Addresses computational challenges in best subset selection
Proposes suboptimal algorithms for high-dimensional data
Compares performance of new and existing methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces new suboptimal algorithm for subset selection
Compares with popular suboptimal algorithms extensively
Focuses on high-dimensional data efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Vikram Singh
University of Central Oklahoma, Edmond, OK
M
Min Sun
The University of Alabama, Tuscaloosa, AL