🤖 AI Summary
This study addresses the significant performance degradation of vision-language models in high-resolution and detail-dense scenarios. To investigate this, we construct a controlled evaluation framework that decouples resolution from task difficulty and propose two novel metrics, AUSC and PVS, to quantify scaling robustness. Through a systematic study integrating semantics-preserving transformations, cross-architecture benchmarking, and fine-grained attribution analysis, this work reveals three primary failure modes: downsampling distortion, tokenization artifacts, and attention dilution. Furthermore, it elucidates the degradation patterns of state-of-the-art models under high-resolution conditions, providing both theoretical foundations and practical guidance for architectural optimization and data augmentation strategies.
📝 Abstract
Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high resolutions is lacking. We introduce a controlled evaluation framework that disentangles resolution-related performance degradation from task difficulty through semantics-preserving transformations. We propose two simple metrics: Area Under the Scaling Curve (AUSC), which quantifies scaling robustness independent of baseline accuracy, and Prediction Variance Score (PVS), which measures resolution-induced prediction instability. Through comprehensive experiments across 5 model families and 5 benchmarks, we identify three primary failure modes: (1) information loss from downsampling at vision token limits, (2) tokenization artifacts from patch boundary shifts and positional encoding fragility under non-standard aspect ratios, and (3) attention dilution as token counts increase. Our analysis reveals that even state-of-the-art models suffer from performance drops when processing high-resolution images, with degradation patterns varying systematically by architectural family. We provide actionable insights for model architecture design and data augmentation strategies to mitigate these limitations.