Unsupervised Long-Tailed Adaptation of Vision-Language Models
This study addresses the severe performance degradation on head classes in vision-language models (VLMs) during unsupervised long-tailed adaptation caused by distribution mismatch. For the first time, this work formally defines this task and elucidates its underlying impairment mechanisms. To overcome these challenges, we propose MARS, which innovatively leverages a zero-shot VLM as a fixed reference. By integrating margin-preserving alignment with perception-gap self-refinement strategies, MARS effectively mitigates pseudo-label bias, corrects inherent model biases, and refines predictions for tail and easily confused classes. Extensive experiments across nine benchmark datasets demonstrate that MARS achieves an average accuracy improvement of 4.71%, significantly outperforming existing state-of-the-art methods.