Institution profile

Shandong Artificial Intelligence Institute

Academic institutionasia · cn
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models

Oct 07, 2026

This study addresses the degradation of generalization capability and representation inconsistency in prompt learning when adapting vision-language models to novel tasks. To this end, we propose DiscoVL, a framework that decouples cross-modal representation learning to eliminate modality interference. Specifically, it introduces a multi-branch low-rank residual aligner to facilitate bidirectional feedback and structured cross-modal alignment, alongside an orthogonal adversarial regularization term that effectively mitigates centroid collapse inherent in triplet loss optimization. Extensive evaluations across 15 benchmarks demonstrate that DiscoVL significantly outperforms state-of-the-art methods in base-to-novel class generalization, cross-dataset transfer, and few-shot learning scenarios. These results validate the efficacy of decoupled representations and orthogonal constraints in enhancing the generalization performance of vision-language models.

0 citationsRead paper
Recent publications

Latest Papers

DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models

Oct 07, 2026

This study addresses the degradation of generalization capability and representation inconsistency in prompt learning when adapting vision-language models to novel tasks. To this end, we propose DiscoVL, a framework that decouples cross-modal representation learning to eliminate modality interference. Specifically, it introduces a multi-branch low-rank residual aligner to facilitate bidirectional feedback and structured cross-modal alignment, alongside an orthogonal adversarial regularization term that effectively mitigates centroid collapse inherent in triplet loss optimization. Extensive evaluations across 15 benchmarks demonstrate that DiscoVL significantly outperforms state-of-the-art methods in base-to-novel class generalization, cross-dataset transfer, and few-shot learning scenarios. These results validate the efficacy of decoupled representations and orthogonal constraints in enhancing the generalization performance of vision-language models.

0 citationsRead paper