Pay More Attention To Text In High-Resolution MLLMs

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文针对高分辨率多模态语言模型中的文本指导问题,提出EviSpec方法以明确视觉证据需求,提高模型性能。
📝 Abstract
Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.
Problem

Research questions and friction points this paper is trying to address.

high-resolution MLLMs
linguistic bottleneck
visual evidence
evidence specification
localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

EviSpec
linguistic bottleneck
evidence specification
high-resolution MLLMs
visual search
Z
Zhongkuan Mao
National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University
W
Wenzhuo Zhao
College of Computer Science, Sichuan University
Xianjie Liu
Xianjie Liu
Master Candidate in Sichuan University
computer vision
Y
Yidong Wang
Peking University
Z
Zhao Gao
College of Computer Science, Sichuan University
R
Ronghao Xian
College of Computer Science, Sichuan University
Y
Yao Jiang
College of Computer Science, Sichuan University
Y
Yi Zhang
X-Humanoid
Liangjian Wen
Liangjian Wen
Southwestern University of Finance and Economics Chengdu, China
Keren Fu
Keren Fu
Sichuan University, College of Computer Science
computer visionimage processingmachine learning