🤖 AI Summary
This study addresses the challenge of model selection under domain shift in the absence of target labels. We propose an unsupervised selection strategy that anchors on a teacher model, revealing that confidence estimators suffer from ranking failure within overfitting collapse regions. Accordingly, we establish a stable selection mechanism based on relative anchoring to the teacher and provide theoretical analysis via quadratic identities. Combined with post-training quantization and statistical bias correction techniques, our approach is evaluated across 134 candidate families spanning CNNs, Vision Transformers, and unsupervised domain adaptation scenarios. Experimental results demonstrate that the proposed method significantly reduces average selection regret using only minimal labeled data, outperforming conventional baselines.
📝 Abstract
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. Established estimators divide sharply: in the overconfident-collapse regime of the CNN families, confidence-based estimators order the family close to backwards, and the diagnostics that identify it need the labels the setting denies, while output-distribution estimators match the teacher-relative anchor and on one architecture beat it. Distortion is nonetheless stable, so a supervised term can move selection away from it. Combining the two, we give exact quadratic identities for a canonical quadratic analogue of the family. We also show that under symmetric corruption the label-dependent part of a criterion linear in the label indicator is multiplied by one common factor whenever its coefficient sums are candidate-invariant, a class holding teacher contrasts and accuracy but not cross-entropy. These characterize the score's components without bounding selection regret. Across one hundred and thirty-four candidate families, one per independently trained convolutional or Vision Transformer teacher, anchoring reduces mean regret at the smallest label budget in every setting, an advantage that fades beyond twenty-five labels.