Adaptive Visual Token Reduction for Accelerated Image Understanding
This study addresses the high computational overhead and loss of spatial structural information in large vision-language models when processing high-resolution images by proposing the ReFIT framework. This framework introduces a pioneering adaptive window reshaping mechanism that overcomes the limitations of conventional fixed cropping. Through Correlation-guided Window Reshaping (RWR) and Instruction-guided Token Refinement (ITR), it achieves dynamic optimization and efficient reduction of visual tokens, effectively preserving critical spatial structural features such as text. Experiments across four VQA benchmarks demonstrate that the proposed method significantly improves answer accuracy while reducing computational costs, validating the effectiveness of regional localization and information deduplication.