GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of vision-language models (VLMs) to over-rely on textual priors while neglecting visual information, which leads to inadequate visual grounding capabilities. To mitigate this issue, we propose GeoPID, a training-free framework that, for the first time, decomposes multimodal representations into redundant, unique, and synergistic components based on subspace geometric relationships. By applying directional intervention during inference to amplify the visually unique components, GeoPID achieves visual enhancement without requiring parameter updates. Extensive experiments across 22 VLMs and 14 benchmarks demonstrate an average relative accuracy improvement of 7.63%, validating the strong generalizability and effectiveness of the proposed approach.
📝 Abstract
While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Information Under-use
Textual Over-reliance
Visual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Training-free Framework
Geometric Decomposition
Representation Subspaces
Visual Grounding
🔎 Similar Papers
No similar papers found.