Score
Designs and analyzes Bayesian variable-selection models that use spike-and-slab priors—mixtures of a point mass at zero (the spike) and a diffuse continuous component (the slab)—to induce sparsity and perform joint inference on parameter values and binary inclusion indicators. Builds hierarchical formulations (including Kuo–Mallick-style parameterizations), specifies prior hyperparameters, and implements computational strategies (e.g., MCMC, marginalization, or variational approximations) to estimate posterior inclusion probabilities, selected model structure, and parameter posteriors.
This study addresses the challenge of simultaneously achieving variable fusion and selection in regression modeling by proposing a novel approach within the Bayesian model averaging framework. The method introduces latent variables to construct a discrete model space, thereby unifying the clustering of covariates—assigning them identical coefficients—and the elimination of irrelevant predictors. It innovatively employs a non-local prior tailored for fusion tasks as the slab component within a spike-and-slab architecture, coupled with an efficient Gibbs sampling scheme. Both theoretical analysis and empirical experiments demonstrate that the proposed method achieves superior performance in terms of fusion accuracy, variable selection consistency, and computational efficiency.
Efficient posterior sampling under spike-and-slab priors remains challenging for high-dimensional sparse linear regression, especially when signal-to-noise ratio (SNR) is arbitrary and the number of measurements $n$ is sublinear in dimension $d$. Method: We propose the first provably efficient sampling algorithm that operates under arbitrary SNR and requires only $n gg k^3 cdot mathrm{polylog}(d)$ measurements. Our approach leverages restricted isometry property (RIP) matrix analysis to design both a polynomial-time Markov chain Monte Carlo (MCMC) sampler and a near-linear stochastic gradient Langevin dynamics (SGLD) variant, compatible with Gaussian and Laplace diffuse components. Contribution/Results: Under $k$-sparsity, our algorithm achieves $mathcal{O}(nd)$ time complexity with rigorous statistical convergence guarantees. It eliminates the need for strong SNR assumptions or $Omega(d)$ measurements required by prior methods, thereby substantially broadening the applicability of Bayesian sparse inference in high dimensions.
This work proposes the first fully Bayesian hierarchical approach for covariate selection in generalized linear models that simultaneously achieves full conjugacy, posterior consistency, and broad applicability across exponential family distributions. By introducing binary inclusion indicators to explicitly model whether each covariate enters the linear predictor, the method unifies variable selection and parameter estimation within a single coherent framework, effectively accounting for model uncertainty. Built upon conjugate priors, the approach enables efficient Gibbs sampling and is accompanied by an R package for practical implementation. Theoretical analysis establishes posterior consistency for both the inclusion indicators and the active regression coefficients. Extensive experiments on synthetic and real-world datasets demonstrate superior performance in terms of predictive accuracy and statistical inference.
This paper addresses Bayesian group-sparse variable selection in high-dimensional generalized linear models (GLMs), unifying logistic, Poisson, negative binomial, and Gaussian regression under canonical or non-canonical link functions. We propose a Bayesian group-regularization framework based on continuous spike-and-slab priors. For the first time, we establish that its maximum a posteriori (MAP) estimator achieves the same minimax L₂ convergence rate as the full posterior distribution, while the posterior contraction rate strictly dominates that of group Lasso. Computationally, the method integrates an EM algorithm with MCMC sampling to ensure both feasibility and theoretical rigor. Extensive simulations and real-data analysis on HIV drug resistance prediction demonstrate that the proposed approach significantly enhances robustness, statistical accuracy, and interpretability in identifying high-dimensional protein sequence features.
This paper addresses group-sparse regression by proposing GSVB, a scalable variational Bayesian framework for generalized linear models with Gaussian, binomial, and Poisson responses. GSVB employs a group-wise spike-and-slab prior and coordinate-ascent variational inference (CAVI). It establishes, for the first time, asymptotic posterior contraction guarantees for variational approximations under group sparsity. Compared to existing maximum-a-posteriori (MAP) methods, GSVB achieves substantially improved variable selection accuracy and more reliable uncertainty quantification. Relative to Markov chain Monte Carlo (MCMC), it delivers orders-of-magnitude speedup (e.g., 10–50× faster) while matching or exceeding MCMC in predictive and inferential performance. Extensive experiments—including simulations and three real-world datasets—demonstrate state-of-the-art results. GSVB thus bridges statistical rigor and computational scalability, enabling robust, efficient, and interpretable group-sparse learning at scale.
This work addresses the challenges in high-dimensional Bayesian regression, where conventional MCMC methods often get trapped in local modes and maximum a posteriori (MAP) estimation fails to quantify uncertainty. To overcome these limitations, the authors propose a hybrid approach that integrates deterministic optimization with stochastic sampling. Specifically, under a heavy-tailed hyperbolic error model, they first employ a two-stage ECM algorithm to efficiently perform variable selection and substantially reduce the model space. Subsequently, Gibbs sampling is conducted within the high posterior probability subspace to enable full posterior inference, complemented by Bayesian model averaging. The proposed method effectively balances computational efficiency, variable selection accuracy, and robust uncertainty quantification. Empirical evaluations on both simulated and real-world datasets demonstrate its superior performance over current state-of-the-art methods.
This work develops a Bayesian asymptotic theory for group-sparse structures in generalized linear models using spike-and-slab priors. By introducing a support-dependent likelihood condition and sparse local asymptotic normality, and leveraging Laplace approximations centered at a pseudo-true parameter, the authors establish an oracle-type Bernstein–von Mises theorem for the fractional posterior. Under conditions on prior concentration, support penalization, recovery geometry, and beta-min separation, they prove that the posterior contracts at the optimal rate, exactly recovers the true support, and converges to the oracle Gaussian distribution. The methodology is validated across multiple regression settings, including Gaussian, logistic, Poisson, probit, gamma, and negative binomial log-link models.
This work addresses the lack of effective methods in existing SLOPE models to simultaneously achieve high predictive performance and rigorous false discovery rate (FDR) control under general design matrices. The authors propose Bayesian Group SLOPE (BGSLOPE) and Bayesian Sparse Group SLOPE (BSGS), which, for the first time, embed the SLOPE penalty within a continuous spike-and-slab Bayesian framework. To restore FDR control in non-orthogonal designs, they introduce a two-step orthogonalization (TSO) strategy. This approach unifies strong statistical power with strict FDR guarantees and enables robust model selection through Bayesian inference and cross-validation. Empirical results on both synthetic and real-world datasets demonstrate that the proposed methods consistently control FDR, substantially improve detection power, and outperform existing approaches in predictive accuracy.