🤖 AI Summary
This work investigates the asymptotic generalization performance of overparameterized linear multiclass classifiers under a Gaussian covariate bilevel model, where sample size, feature dimension, and number of classes all diverge simultaneously. Addressing Subramanian et al. (2022)’s conjecture on the suboptimality of interpolating classifiers, we provide the first rigorous theoretical confirmation: in a specific high-dimensional regime, the minimum-norm interpolating classifier exhibits an asymptotically higher misclassification rate than non-interpolating counterparts—demonstrating a phase transition in relative performance. Technically, we introduce a novel variant of the Hanson–Wright inequality tailored to sparse label settings, and integrate it with asymptotic random matrix theory and information-theoretic strong converse bounds to derive tight upper and lower bounds on generalization error—establishing that the misclassification rate must converge almost surely to either zero or one. Our framework further extends successfully to multi-label classification.
📝 Abstract
We study the asymptotic generalization of an overparameterized linear model for multiclass classification under the Gaussian covariates bi-level model introduced in Subramanian et al.~'22, where the number of data points, features, and classes all grow together. We fully resolve the conjecture posed in Subramanian et al.~'22, matching the predicted regimes for generalization. Furthermore, our new lower bounds are akin to an information-theoretic strong converse: they establish that the misclassification rate goes to 0 or 1 asymptotically. One surprising consequence of our tight results is that the min-norm interpolating classifier can be asymptotically suboptimal relative to noninterpolating classifiers in the regime where the min-norm interpolating regressor is known to be optimal. The key to our tight analysis is a new variant of the Hanson-Wright inequality which is broadly useful for multiclass problems with sparse labels. As an application, we show that the same type of analysis can be used to analyze the related multilabel classification problem under the same bi-level ensemble.