🤖 AI Summary
This study addresses the phenomenon that covariate shift can induce reliability drift in classifiers even when their confidence distributions remain invariant. The authors investigate worst-case drift under a χ²-divergence budget constraint, constructing vulnerability curves and estimating their lower bounds. They demonstrate that calibration residuals and group variances cannot solely determine vulnerability, proposing instead a restricted estimation method based on finite readout layers. Robust inference is achieved by integrating confidence binning, role-separated labels, and out-of-sample lower bound estimation techniques. Empirical validation on ImageNet confirms positive lower bounds across multiple classifiers, with optimized reweighted actual drift closely matching estimated curves. This work provides both theoretical guarantees and practical tools for evaluating model reliability under distribution shifts.
📝 Abstract
A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliability relation under covariate shifts that preserve the distribution of the confidence score, constraining the reweighting within each confidence level by a $χ^2$ budget; the resulting worst case, as a function of the budget, is a fragility profile. On an interval of budgets that can be computed from the source distribution, the profile equals exactly the square root of the budget times the within-level variance of the correctness propensity -- the grouping-loss term of calibration-refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within-level law; consequently, calibration residual and grouping variance do not determine fragility in general, though they do when labels and predictions are deterministic. Since the propensity is not observed, we restrict reweightings to a learned finite readout within confidence bins, bound the part the restriction misses by the grouping variance remaining inside readout cells, estimate the restricted profile with role-separated labels, and provide a separate split-sample lower confidence bound. On ImageNet this bound is positive in both splits for four of six primary classifiers and nine of twelve additional ones as released, and for three of eighteen after temperature scaling. Held-out drift under optimised reweightings fitted without evaluation labels tracks the estimated profile; an exploratory label-permutation diagnostic yields near-zero agreement for this statistic while largely reproducing the correlation observed for unsigned random reweightings.