🤖 AI Summary
Modeling high-dimensional over-dispersed count data poses dual challenges in computation and variable selection, as traditional Bayesian negative binomial regression relies on Markov chain Monte Carlo (MCMC) methods that do not scale well. This work proposes the first efficient sparse negative binomial regression framework by integrating variational Bayesian inference with both the horseshoe prior and continuous shrinkage priors. The resulting approach achieves estimation accuracy and variable selection performance comparable to MCMC while accelerating computation by over two orders of magnitude. Moreover, it demonstrates robustness across both Poisson and over-dispersed settings, making it a compelling default method for high-dimensional count data modeling.
📝 Abstract
Count data with overdispersion and high-dimensional predictors pose significant challenges in modern applications. While negative binomial regression offers a flexible modeling framework, existing Bayesian approaches rely on computationally expensive MCMC methods that become impractical in high-dimensional settings. This paper develops a variational Bayesian framework for sparse negative binomial regression using horseshoe and continuous shrinkage priors. Our proposed methods achieve estimation accuracy and variable selection performance comparable to MCMC benchmarks while requiring less than 1\% of the computation time. Extensive simulations demonstrate that the negative binomial specification is essential for overdispersed data, as Poisson-based approaches exhibit substantial performance degradation under overdispersion. Conversely, our methods remain robust when the data are Poisson, making them a safer default choice. Applications to real benchmark datasets further confirm the practical utility of our approach. The proposed framework provides a computationally efficient and reliable tool for sparse count regression in high-dimensional settings.