High-Dimensional Statistical Inference for Sparse Support Vector Machines

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of statistical inference in high-dimensional sparse support vector machines, where the non-smoothness of the hinge loss complicates analysis. To overcome this obstacle, we propose an inference framework that integrates a symmetrized data representation, linear programming, and duality theory to achieve debiased estimation and establish coordinate-wise asymptotic normality. This approach enables the construction of valid confidence intervals and hypothesis tests, facilitating variable selection with controlled false discovery rates. Simulation studies confirm the method’s calibration accuracy and statistical power. Furthermore, an application to breast cancer gene expression data demonstrates its capacity to precisely identify significant features, effectively distinguishing true signals from variables spuriously selected by the original model.
πŸ“ Abstract
Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
Problem

Research questions and friction points this paper is trying to address.

High-dimensional inference
Sparse support vector machines
Nonsmooth hinge loss
Debiasing
Variable selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Support Vector Machine
High-Dimensional Inference
Debiased Estimator
Hinge Loss
False Discovery Rate