π€ AI Summary
This study addresses the challenge of statistical inference in high-dimensional sparse support vector machines, where the non-smoothness of the hinge loss complicates analysis. To overcome this obstacle, we propose an inference framework that integrates a symmetrized data representation, linear programming, and duality theory to achieve debiased estimation and establish coordinate-wise asymptotic normality. This approach enables the construction of valid confidence intervals and hypothesis tests, facilitating variable selection with controlled false discovery rates. Simulation studies confirm the methodβs calibration accuracy and statistical power. Furthermore, an application to breast cancer gene expression data demonstrates its capacity to precisely identify significant features, effectively distinguishing true signals from variables spuriously selected by the original model.
π Abstract
Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.