🤖 AI Summary
This study addresses the degradation of generalization capability and representation inconsistency in prompt learning when adapting vision-language models to novel tasks. To this end, we propose DiscoVL, a framework that decouples cross-modal representation learning to eliminate modality interference. Specifically, it introduces a multi-branch low-rank residual aligner to facilitate bidirectional feedback and structured cross-modal alignment, alongside an orthogonal adversarial regularization term that effectively mitigates centroid collapse inherent in triplet loss optimization. Extensive evaluations across 15 benchmarks demonstrate that DiscoVL significantly outperforms state-of-the-art methods in base-to-novel class generalization, cross-dataset transfer, and few-shot learning scenarios. These results validate the efficacy of decoupled representations and orthogonal constraints in enhancing the generalization performance of vision-language models.
📝 Abstract
Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning