🤖 AI Summary
This paper investigates calibration of high-dimensional linear binary classifiers under the proportional asymptotic regime where the feature dimension $p$ and sample size $n$ grow at comparable rates ($p/n o gamma in (0,infty)$). We address miscalibration arising from bias in weight estimation inherent to conventional linear predictors. To resolve this, we propose *angular calibration*: a novel procedure that leverages a consistent estimator of the angle between the estimated and true weight vectors to interpolate between the learned classifier and a random-guess classifier, thereby achieving exact calibration. We establish that, in the high-dimensional limit, this method simultaneously satisfies both calibration and uniqueness-optimality under Bregman divergence—marking the first result to guarantee both properties jointly. Furthermore, we prove that Platt scaling converges to this optimal angular calibration solution in the proportional regime, providing a theoretical foundation for its empirical success.
📝 Abstract
We study the fundamental problem of calibrating a linear binary classifier of the form $sigma(hat{w}^ op x)$, where the feature vector $x$ is Gaussian, $sigma$ is a link function, and $hat{w}$ is an estimator of the true linear weight $w^star$. By interpolating with a noninformative $ extit{chance classifier}$, we construct a well-calibrated predictor whose interpolation weight depends on the angle $angle(hat{w}, w_star)$ between the estimator $hat{w}$ and the true linear weight $w_star$. We establish that this angular calibration approach is provably well-calibrated in a high-dimensional regime where the number of samples and features both diverge, at a comparable rate. The angle $angle(hat{w}, w_star)$ can be consistently estimated. Furthermore, the resulting predictor is uniquely $ extit{Bregman-optimal}$, minimizing the Bregman divergence to the true label distribution within a suitable class of calibrated predictors. Our work is the first to provide a calibration strategy that satisfies both calibration and optimality properties provably in high dimensions. Additionally, we identify conditions under which a classical Platt-scaling predictor converges to our Bregman-optimal calibrated solution. Thus, Platt-scaling also inherits these desirable properties provably in high dimensions.