Can't Find Waldo: Evaluating VLMs'Sensitivity to Image Resolution and Detail Level

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation of vision-language models in high-resolution and detail-dense scenarios. To investigate this, we construct a controlled evaluation framework that decouples resolution from task difficulty and propose two novel metrics, AUSC and PVS, to quantify scaling robustness. Through a systematic study integrating semantics-preserving transformations, cross-architecture benchmarking, and fine-grained attribution analysis, this work reveals three primary failure modes: downsampling distortion, tokenization artifacts, and attention dilution. Furthermore, it elucidates the degradation patterns of state-of-the-art models under high-resolution conditions, providing both theoretical foundations and practical guidance for architectural optimization and data augmentation strategies.
📝 Abstract
Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high resolutions is lacking. We introduce a controlled evaluation framework that disentangles resolution-related performance degradation from task difficulty through semantics-preserving transformations. We propose two simple metrics: Area Under the Scaling Curve (AUSC), which quantifies scaling robustness independent of baseline accuracy, and Prediction Variance Score (PVS), which measures resolution-induced prediction instability. Through comprehensive experiments across 5 model families and 5 benchmarks, we identify three primary failure modes: (1) information loss from downsampling at vision token limits, (2) tokenization artifacts from patch boundary shifts and positional encoding fragility under non-standard aspect ratios, and (3) attention dilution as token counts increase. Our analysis reveals that even state-of-the-art models suffer from performance drops when processing high-resolution images, with degradation patterns varying systematically by architectural family. We provide actionable insights for model architecture design and data augmentation strategies to mitigate these limitations.
Problem

Research questions and friction points this paper is trying to address.

Visual Language Models
Image Resolution
Performance Degradation
Failure Modes
High-resolution Images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Language Models
High-Resolution Evaluation
Semantics-Preserving Transformations
Scaling Robustness Metrics
Failure Modes Analysis