Random Subset Averaging

📅 2025-12-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address prediction instability and inaccuracy under high-dimensional, strongly correlated covariates, this paper proposes the Random Subset Averaging (RSA) ensemble method: base models are constructed via binomial random sampling, and adaptively fused through two rounds of data-driven weighting—structurally analogous to a two-layer neural network. Theoretical contributions include: (i) the first asymptotic optimality guarantee for two-stage adaptive weighting; (ii) allowance for fully data-dependent weights in the first stage; and (iii) tighter finite-sample risk bounds under orthogonal design. RSA integrates cross-validation-based hyperparameter tuning with dual-stage weighting, ensuring both computational feasibility and statistical rigor. Empirical results demonstrate that RSA significantly outperforms LASSO, random forests, and state-of-the-art ensemble methods across diverse sparse and highly correlated settings. Its practical robustness and superiority are further validated in financial return forecasting.

Technology Category

Machine Learning: Ensemble MethodsReasoning under Uncertainty: Stochastic OptimizationIntelligent Robots: State Estimation

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
We propose a new ensemble prediction method, Random Subset Averaging (RSA), tailored for settings with many covariates, particularly in the presence of strong correlations. RSA constructs candidate models via binomial random subset strategy and aggregates their predictions through a two-round weighting scheme, resulting in a structure analogous to a two-layer neural network. All tuning parameters are selected via cross-validation, requiring no prior knowledge of covariate relevance. We establish the asymptotic optimality of RSA under general conditions, allowing the first-round weights to be data-dependent, and demonstrate that RSA achieves a lower finite-sample risk bound under orthogonal design. Simulation studies demonstrate that RSA consistently delivers superior and stable predictive performance across a wide range of sample sizes, dimensional settings, sparsity levels and correlation structures, outperforming conventional model selection and ensemble learning methods. An empirical application to financial return forecasting further illustrates its practical utility.
Problem

Research questions and friction points this paper is trying to address.

Develops ensemble method for high-dimensional correlated data
Optimizes predictions via two-round weighting and cross-validation
Outperforms existing methods in simulations and financial forecasting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ensemble method uses binomial random subset strategy
Two-round weighting scheme aggregates predictions like neural network
Cross-validation selects all tuning parameters without prior knowledge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wenhao Cui
School of Economics and Management, Beihang University
J
Jie Hu
Yau Mathematical Sciences Center, Tsinghua University