🤖 AI Summary
This study addresses the severe performance degradation on head classes in vision-language models (VLMs) during unsupervised long-tailed adaptation caused by distribution mismatch. For the first time, this work formally defines this task and elucidates its underlying impairment mechanisms. To overcome these challenges, we propose MARS, which innovatively leverages a zero-shot VLM as a fixed reference. By integrating margin-preserving alignment with perception-gap self-refinement strategies, MARS effectively mitigates pseudo-label bias, corrects inherent model biases, and refines predictions for tail and easily confused classes. Extensive experiments across nine benchmark datasets demonstrate that MARS achieves an average accuracy improvement of 4.71%, significantly outperforming existing state-of-the-art methods.
📝 Abstract
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.