🤖 AI Summary
Vision-language models such as CLIP are susceptible to significant classification degradation under distribution shifts due to embedding drift. To address this, this work proposes DRC, a training-free calibration method that leverages unlabeled target-domain images to mitigate cross-domain embedding shifts without fine-tuning. Specifically, DRC eliminates hard-clustering bias through Gaussian mixture re-centering and confidence-weighted prior correction, while achieving feature alignment via posterior-weighted averaging and log-prior adjustment. Experimental results demonstrate that DRC substantially outperforms the zero-shot CLIP baseline in cross-domain accuracy across diverse scenarios, yielding improvements of up to 5.07 percentage points.
📝 Abstract
Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of unlabeled target images. DRC fits a Gaussian mixture once and subtracts from each embedding a posterior-weighted average of component means. It then removes residual class preference with a log-prior correction, estimating the prior from confidence-weighted predictions. Among compared methods, DRC achieves the highest average accuracy on cross-domain datasets, exceeding zero-shot CLIP by 4.13 and 5.07 points with ViT-B/16 and ResNet-50, with gains over CLIP also holding under ImageNet distribution shifts.