Seeing and Solving Are Not Enough for Vision-Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the compositional failures in vision-language models (VLMs), where incorrect responses frequently arise from the breakdown between visual extraction and reasoning capabilities. To tackle this issue, this work presents the first quantitative analysis of such phenomena and proposes a state-realization fine-tuning strategy. Leveraging LoRA adapters within an autoregressive paradigm, the approach compels the model to explicitly generate precise, scoreable task states prior to producing final answers, thereby bridging the gap between perception and reasoning. Experimental results demonstrate that the proposed method improves accuracy by 1.7 to 14.1 percentage points and successfully resolves 92.5% to 98.1% of compositional failure cases. Furthermore, it exhibits strong cross-task generalization capabilities, highlighting its effectiveness as a robust solution for enhancing the compositional reliability of VLMs.
📝 Abstract
Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Composition Failures
Multimodal Reasoning
Visual Extraction
Problem Solving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Composition Failures
State Realization Tuning
LoRA Adapters
Task State
🔎 Similar Papers
Z
Ziheng Wang
Sun Yat-sen University
M
Mingxuan Xie
Zhejiang University
Yilin Liu
Yilin Liu
Google
AI/MLWearable devicesMotion sensingHealthcare AI
D
Dayan Wu
Institute of Information Engineering
Y
Yang Li
Hunan University
P
Pengwen Dai
Sun Yat-sen University