ScreenHaystack: Finding Blind Zones in GUI Grounding

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "blind spot" problem in GUI grounding models, where localization accuracy degrades sharply in specific screen regions due to spatially imbalanced training data. We present the first systematic quantification of this phenomenon by proposing a dynamic benchmarking framework that evaluates spatial reliability through target icon repositioning, complemented by controlled synthetic experiments confirming the correlation between blind spots and data distribution. Building on these findings, we introduce a blind-spot-aware supervised data augmentation strategy. Experiments reveal significant spatial blind spots in mainstream models such as Qwen3-VL. When fine-tuned with our proposed augmentation, models achieve substantially higher accuracy on ScreenSpot-Pro compared to both original and randomly augmented baselines, while demonstrating strong cross-model transferability.
📝 Abstract
We introduce ScreenHaystack, a dynamic needle-in-a-haystack benchmark for evaluating spatial reliability in GUI grounding. Instead of testing each target at a fixed position, ScreenHaystack systematically relocates controlled target icons across high-resolution GUI backgrounds and measures whether models can localize them consistently. Using this benchmark, we find that leading GUI grounding models, including Qwen3-VL, UI-TARS, GTA, and UI-Venus, exhibit blind zones: spatial regions where grounding accuracy drops sharply despite fixed target appearance and instruction. These blind zones transfer to unseen ScreenSpot-Pro examples: targets inside blind zones are consistently harder to ground, with Qwen3-VL-8B dropping by 16.1 percentage points, and controlled relocation shows that moving targets into blind zones decreases accuracy while moving them out improves accuracy. We further show through controlled synthetic experiments that uneven spatial coverage in training data can induce such blind zones. Therefore, we propose a simple strategy, blind-zone-oriented augmentation, which adds supervision in blind zones and improves ScreenSpot-Pro accuracy over both original and randomly augmented Click-100k fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

GUI grounding
blind zones
spatial reliability
needle-in-a-haystack benchmark
training data coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

GUI Grounding
Blind Zones
Needle-in-a-Haystack Benchmark
Spatial Reliability
Data Augmentation