Score
Design, implement, and evaluate algorithms and workflows that identify, rank, and select compact, informative, and often sparse subsets of input variables (features) for predictive or analysis models. This includes methods that enforce hierarchical or temporal (lag) inclusion, improve robustness and reproducibility via adversarial or stability-driven criteria, perform Bayesian variable selection using MCMC/Gibbs or stochastic-search strategies, and use embedded, label-assisted, greedy, or spectral/statistical ranking heuristics to produce interpretable feature supports.
This paper addresses linear time-series regression models featuring autocorrelated errors and lagged covariates, proposing the first Bayesian framework for *joint selection* of important covariates and the autoregressive order of the error process. Methodologically, we develop a hierarchical spike-and-slab prior that enables simultaneous Bayesian variable selection for both regressors and error lags—novel in the literature—and design a two-stage MCMC algorithm to enhance computational efficiency and selection accuracy in high-dimensional settings. We establish theoretical high-dimensional consistency under mild regularity conditions. Empirical evaluations on groundwater depth forecasting and S&P 500 log-return modeling demonstrate substantial reductions in mean squared prediction error (MSPE), improved identification of the true model, and enhanced predictive robustness. The method is particularly effective in domains with strong temporal dependence, including finance, hydrology, and meteorology.
This work proposes the first fully Bayesian hierarchical approach for covariate selection in generalized linear models that simultaneously achieves full conjugacy, posterior consistency, and broad applicability across exponential family distributions. By introducing binary inclusion indicators to explicitly model whether each covariate enters the linear predictor, the method unifies variable selection and parameter estimation within a single coherent framework, effectively accounting for model uncertainty. Built upon conjugate priors, the approach enables efficient Gibbs sampling and is accompanied by an R package for practical implementation. Theoretical analysis establishes posterior consistency for both the inclusion indicators and the active regression coefficients. Extensive experiments on synthetic and real-world datasets demonstrate superior performance in terms of predictive accuracy and statistical inference.
This work addresses the vulnerability of predictive models in high-stakes domains—such as healthcare—to strategic manipulation of input features, a challenge inadequately mitigated by existing coarse-grained feature filtering approaches. The authors propose a novel method that jointly optimizes feature selection and the regularization strength of ridge regression to enhance robustness against such strategic behavior. By integrating game-theoretic modeling with feature selection theory, they demonstrate that naively discarding features based solely on manipulability is often suboptimal and instead provide a refined characterization of the performance of feature subsets under strategic perturbations. Empirical validation on real-world healthcare payment data confirms the efficacy of the proposed algorithm, offering a principled and practical framework for designing coarse-grained policies resilient to strategic manipulation.
This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.
Machine learning model selection lacks formalized methodologies, making it difficult to systematically characterize contextual factors—such as data characteristics and prediction tasks—and their interactions, resulting in opaque, non-adaptive decisions. This paper introduces, for the first time, software product line (SPL) principles into ML model selection, proposing a variability-aware algorithm selection framework. It constructs a configurable feature model that explicitly captures commonalities and variabilities among contextual factors—including dataset size, feature dimensionality, and task type—as well as their logical dependencies. By integrating scikit-learn’s heuristic rules with an instantiation framework, the approach enables interpretable, adaptive, and transparent model recommendations. An empirical case study demonstrates that the method significantly outperforms existing strategies in accuracy, interpretability, and contextual adaptability.
This study addresses the challenge that traditional Bayesian variable selection methods struggle to accurately quantify uncertainty under model misspecification, thereby compromising selection performance. The authors propose a quasi-posterior–based variable selection approach that requires only the specification of mean and variance functions, eliminating the need for a fully specified likelihood. This method combines robustness with the advantages of Bayesian inference. For the first time, the quasi-posterior framework is systematically introduced into variable selection, leveraging Laplace approximation to efficiently compute quasi-marginal likelihoods. By avoiding full model specification while preserving desirable Bayesian properties, the proposed method achieves substantially improved selection accuracy under complex data-generating mechanisms—such as heavy-tailed errors and overdispersed count outcomes—and demonstrates strong empirical performance on real-world datasets from social sciences and genomics.
Feature selection in high-dimensional regression is highly susceptible to the dual perturbations of sampling variability and measurement error in the design matrix. This work proposes a perturbation-and-aggregation framework that injects controlled additive noise into subsampled data and evaluates feature selection frequencies across multiple noise levels to construct stability paths, thereby identifying features robust to both types of perturbations. The method uniquely models sampling randomness and design noise simultaneously while preserving full-sample utilization and maintaining compatibility with base selectors such as Lasso. Theoretical analysis establishes model selection consistency under small perturbations, and empirical results demonstrate that the approach significantly outperforms existing methods on both synthetic and real-world datasets, exhibiting superior robustness.
Traditional Bayesian modeling relies on model selection to balance complexity and generalization, yet this approach often compromises predictive performance in small-sample settings. This work proposes a “predictive consistency prior” that maintains stability in the prior predictive distribution as model complexity increases, thereby circumventing explicit model selection. By shifting the modeling focus from parameter sparsity to constructing reasonable and stable priors in predictive space, the method reveals that the perceived necessity of model selection fundamentally stems from inadequate prior specification. The authors implement this prior in Bayesian linear and logistic regression, forward variable selection, and nonlinear models, demonstrating through numerical experiments that flexible models equipped with the predictive consistency prior match or even outperform carefully selected simpler models in out-of-sample prediction across a range of tasks.
This work addresses the challenges in high-dimensional Bayesian regression, where conventional MCMC methods often get trapped in local modes and maximum a posteriori (MAP) estimation fails to quantify uncertainty. To overcome these limitations, the authors propose a hybrid approach that integrates deterministic optimization with stochastic sampling. Specifically, under a heavy-tailed hyperbolic error model, they first employ a two-stage ECM algorithm to efficiently perform variable selection and substantially reduce the model space. Subsequently, Gibbs sampling is conducted within the high posterior probability subspace to enable full posterior inference, complemented by Bayesian model averaging. The proposed method effectively balances computational efficiency, variable selection accuracy, and robust uncertainty quantification. Empirical evaluations on both simulated and real-world datasets demonstrate its superior performance over current state-of-the-art methods.