Robust Multi-Model Subset Selection

📅 2023-11-22
📈 Citations: 0
Influential: 0
📄 PDF

career value

211K/year
🤖 AI Summary
Addressing the challenge of robust prediction in high-dimensional data corrupted by outliers, this paper proposes a sparse, diverse, and robust ensemble modeling framework. Methodologically, we establish, for the first time, the finite-sample breakdown point theory for ensemble models; design a custom L₀-optimization-driven algorithm that jointly controls sparsity, diversity, and robustness; and integrate robust statistical modeling, ensemble learning, and cross-validation-based hyperparameter tuning. Our key contribution lies in introducing L₀ regularization into robust ensemble learning—enabling synergistic enhancement of outlier identification and outlier resistance. Experiments on contaminated bioinformatics and cheminformatics datasets demonstrate that the proposed method significantly outperforms existing single-model sparse robust approaches, achieving higher predictive accuracy and more reliable outlier detection.
📝 Abstract
Outlying observations can be challenging to handle and adversely affect subsequent analyses, particularly, in complex high-dimensional datasets. Although outliers are not always undesired anomalies in the data and may possess valuable insights, only methods that are robust to outliers are able to accurately identify them and resist their influence. In this paper, we propose a method that generates an ensemble of sparse and diverse predictive models that are resistant to outliers. We show that the ensembles generally outperform single-model sparse and robust methods in high-dimensional prediction tasks. Cross-validation is used to tune model parameters to control levels of sparsity, diversity and resistance to outliers. We establish the finitesample breakdown point of the ensembles and the models that comprise them, and we develop a tailored computing algorithm to learn the ensembles by leveraging recent developments in L0 optimization. Our extensive numerical experiments on synthetic and artificially contaminated real datasets from bioinformatics and cheminformatics demonstrate the competitive advantage of our method over state-of-the-art single-model methods.
Problem

Research questions and friction points this paper is trying to address.

Handling outlying observations in complex data
Generating robust sparse diverse predictive model ensembles
Outperforming single-model methods with outlier resistance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robust Multi-Model Subset Selection method
Generates ensemble of sparse diverse models
Uses l0 optimization tailored computing algorithm
🔎 Similar Papers
No similar papers found.