Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis

๐Ÿ“… 2026-10-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitations of vision-language models in fine-grained perception and the shortcomings of existing alignment methods by proposing a representation alignment framework based on Kernel Canonical Correlation Analysis (KCCA). Methodologically, KCCA is introduced to maximize cross-modal projection correlation, while Karushโ€“Kuhnโ€“Tucker (KKT) conditions are leveraged to circumvent eigendecomposition, thereby enabling end-to-end training. The framework is further extended to joint image-text-video tri-view alignment, achieving unified optimization of multi-view learning and CLIP fine-tuning. Experimental results demonstrate that the proposed method significantly improves ImageNet accuracy to 25.9% on the MMVP-VLM benchmark, outperforming existing approaches while preserving zero-shot retrieval performance.
๐Ÿ“ Abstract
Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Fine-grained Visual Perception
Visual Representation Enhancement
Kernel Canonical Correlation Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kernel Canonical Correlation Analysis
Visual Representation Enhancement
Vision-Language Models
End-to-End Optimization
Multi-view Alignment
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.