DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of generalization capability and representation inconsistency in prompt learning when adapting vision-language models to novel tasks. To this end, we propose DiscoVL, a framework that decouples cross-modal representation learning to eliminate modality interference. Specifically, it introduces a multi-branch low-rank residual aligner to facilitate bidirectional feedback and structured cross-modal alignment, alongside an orthogonal adversarial regularization term that effectively mitigates centroid collapse inherent in triplet loss optimization. Extensive evaluations across 15 benchmarks demonstrate that DiscoVL significantly outperforms state-of-the-art methods in base-to-novel class generalization, cross-dataset transfer, and few-shot learning scenarios. These results validate the efficacy of decoupled representations and orthogonal constraints in enhancing the generalization performance of vision-language models.
📝 Abstract
Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Prompt Learning
Cross-Modal Representation
Generalization
Centroid Collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Disentangled Representation Learning
Orthogonal Adversarial Regularization
Low-Rank Residual Aligner
Cross-Modal Alignment
Prompt Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mengping Dong
Shandong Artificial Intelligence Institute, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China
Jinbao Li
Jinbao Li
Department of Geography, University of Hong Kong
Climate ChangePaleoclimateENSODroughtDendrochronology
F
Fei Li
University of Florida, Florida, USA