High Dimensional Gaussian and Bootstrap Approximations in Generalized Linear Models

📅 2026-01-14
📈 Citations: 0
Influential: 0
📄 PDF

career value

215K/year
🤖 AI Summary
This study addresses distributional approximation and statistical inference for high-dimensional generalized linear models (GLMs) where the parameter dimension $d$ grows with or even far exceeds the sample size $n$. By integrating high-dimensional Gaussian approximation, Bahadur representation, and Nazarov’s isoperimetric inequality, the authors establish, for the first time, optimal Gaussian approximation rates for GLM estimators over convex sets under $d = o(n^{2/5})$ and over Euclidean balls under $d = o(n^{1/2})$. To overcome the well-known difficulty that Lasso estimators cannot simultaneously achieve variable selection consistency and $\sqrt{n}$-consistency, a perturbed bootstrap procedure is proposed, accompanied by non-asymptotic Berry–Esseen-type error bounds. Both theoretical analysis and simulation studies demonstrate that the proposed method substantially improves inferential accuracy in high-dimensional GLMs.

Technology Category

Application Category

📝 Abstract
Generalized Linear Models (GLMs) extend ordinary linear regression by linking the mean of the response variable to covariates through appropriate link functions. This paper investigates the asymptotic behavior of GLM estimators when the parameter dimension $d$ grows with the sample size $n$. In the first part, we establish Gaussian approximation results for the distribution of a properly centered and scaled GLM estimator uniformly over class of convex sets and Euclidean balls. Using high-dimensional results from Fang and Koike (2024) for the leading Bahadur term, bounding remainder terms as in He and Shao (2000), and applying Nazarov's (2003) Gaussian isoperimetric inequality, we show that Gaussian approximation holds when $d = o(n^{2/5})$ for convex sets and $d = o(n^{1/2})$ for Euclidean balls-the best possible rates matching those for high-dimensional sample means. We further extend these results to the bootstrap approximation when the covariance matrix is unknown. In the second part, when $d>>n$, a natural question is to answer whether all covariates are equally important. To answer that, we employ sparsity in GLM through the Lasso estimator. While Lasso is widely used for variable selection, it cannot achieve both Variable Selection Consistency (VSC) and $n^{1/2}$-consistency simultaneously (Lahiri, 2021). Under the regime ensuring VSC, we show that Gaussian approximation for the Lasso estimator fails. To overcome this, we propose a Perturbation Bootstrap (PB) approach and establish a Berry-Esseen type bound for its approximation uniformly over class of convex sets. Simulation studies confirm the strong finite-sample performance of the proposed method.
Problem

Research questions and friction points this paper is trying to address.

High-dimensional inference
Generalized Linear Models
Gaussian approximation
Bootstrap approximation
Lasso estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian approximation
Bootstrap approximation
High-dimensional GLM
Perturbation Bootstrap
Lasso estimator