Unsupervised Long-Tailed Adaptation of Vision-Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe performance degradation on head classes in vision-language models (VLMs) during unsupervised long-tailed adaptation caused by distribution mismatch. For the first time, this work formally defines this task and elucidates its underlying impairment mechanisms. To overcome these challenges, we propose MARS, which innovatively leverages a zero-shot VLM as a fixed reference. By integrating margin-preserving alignment with perception-gap self-refinement strategies, MARS effectively mitigates pseudo-label bias, corrects inherent model biases, and refines predictions for tail and easily confused classes. Extensive experiments across nine benchmark datasets demonstrate that MARS achieves an average accuracy improvement of 4.71%, significantly outperforming existing state-of-the-art methods.
📝 Abstract
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Unsupervised Adaptation
Long-Tailed Distribution
Distribution Mismatch
Pseudo-labels
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Long-Tailed Adaptation
Vision-Language Models
Boundary-Preserving Alignment
Margin-aware Self-Refinement
Distribution Mismatch
🔎 Similar Papers
No similar papers found.
K
Keliang Chen
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China
Y
Yaxin Hou
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China
H
Hui Liu
School of Computing and Information Sciences, Saint Francis University, Hong Kong, China
Y
Yuheng Jia
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China; School of Computing and Information Sciences, Saint Francis University, Hong Kong, China