Scalable Variable Selection under Predictor Dependence with Adaptive Virtual Dummies

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the false discovery rate (FDR) control bias and computational redundancy caused by predictor dependence in high-dimensional variable selection. We propose memory-efficient, adaptive virtual dummy variables based on conditional Gaussian laws to correct competition bias, and design a cross-experiment shared lazy Gram caching mechanism to accelerate forward selection path termination. By integrating the T-Rex selector, conditional sampling, and lazy evaluation algorithms, the proposed method achieves scalable high-dimensional statistical inference. Experimental results demonstrate that it stably controls the FDR at target levels across diverse covariance structures while improving computational speed by over sixfold, thereby supporting large-scale applications involving up to one hundred thousand predictors.
📝 Abstract
Reliable high-dimensional variable selection requires scalable error-controlling methods. The Terminating-Random Experiments (T-Rex) selector estimates the false discovery rate (FDR) by aggregating early-terminated forward-selection paths in which predictors compete with synthetic dummies. We address two remaining challenges: i) predictor dependence can bias dummy-predictor competition; ii) computation is wasted on recomputing terms shared across experiments. Building on memory-efficient virtual dummies, which sequentially sample projections from their exact conditional law, we estimate the conditional Gaussian law of inactive predictors and map draws onto the remaining sphere radius. A shared lazy Gram cache computes the response product and each requested Gram column once across experiments. Simulations across three covariance structures show that uniform spherical dummies can exceed the target FDR, whereas the proposed method empirically controls FDR. Caching yields more than a sixfold speedup, and the method remains feasible with 100000 predictors, where competing FDR-controlling methods become computationally impractical.
Problem

Research questions and friction points this paper is trying to address.

variable selection
false discovery rate
predictor dependence
high-dimensional data
computational scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variable Selection
False Discovery Rate
Virtual Dummies
Predictor Dependence
Gram Cache
🔎 Similar Papers
No similar papers found.