Adaptive Visual Token Reduction for Accelerated Image Understanding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead and loss of spatial structural information in large vision-language models when processing high-resolution images by proposing the ReFIT framework. This framework introduces a pioneering adaptive window reshaping mechanism that overcomes the limitations of conventional fixed cropping. Through Correlation-guided Window Reshaping (RWR) and Instruction-guided Token Refinement (ITR), it achieves dynamic optimization and efficient reduction of visual tokens, effectively preserving critical spatial structural features such as text. Experiments across four VQA benchmarks demonstrate that the proposed method significantly improves answer accuracy while reducing computational costs, validating the effectiveness of regional localization and information deduplication.
📝 Abstract
Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.
Problem

Research questions and friction points this paper is trying to address.

Visual Token Reduction
Large Vision-Language Models
Computational Efficiency
Spatial Structure Preservation
Visual Question Answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Token Reduction
Vision-Language Models
Instruction-Guided
Window Reshaping
Token Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Seyoung Jeong
Jeonbuk National University, Jeonju, Republic of Korea
J
Jong Pil Yun
Korea Institute of Industrial Technology (KITECH), Incheon, Republic of Korea; Chung-Ang University, Seoul, Republic of Korea
Sang Jun Lee
Sang Jun Lee
stradvision
SLAMSensor FusionNeural RenderingSpatial AI