Prior Directions: Why GUI Grounding Gets Locked in the Past

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the susceptibility of vision-language models to outdated linguistic descriptions following scene changes, which induces a “visual locking” phenomenon and leads to erroneous visual grounding. The authors introduce the concept of “prior directions” and, through multi-model comparisons, representational geometry analysis, and controlled intervention experiments, demonstrate that the locking effect arises because prior language inputs steer internal representations along a compact and reproducible set of directions in representation space. Crucially, removing representational components aligned with these prior directions substantially restores the model’s visual grounding capability, whereas ablating orthogonal components has negligible impact, thereby validating the pivotal role of prior directions in mediating this interference.
📝 Abstract
Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the verbalized prior varies. Across models, stronger lock-in accompanies smaller changes in the model representation before the final answer. This reversal suggests that lock-in depends not on how far this representation moves, but on how that movement is organized. In models that are harder to correct, prior-induced changes concentrate along a compact set of directions that repeatedly appear across examples. We call these recurrent axes the Prior Directions. They recur on held-out examples, while a descriptive four-model comparison associates greater concentration with stronger lock-in. Controlled interventions show that removing the component aligned with the Prior Directions restores visual grounding, whereas removing an equally large orthogonal component has little effect. Prior control thus arises when prior-induced changes form a coherent and reusable pattern in the representation used to produce the answer. This account explains why the same prior remains revisable in one model yet becomes dominant in another.
Problem

Research questions and friction points this paper is trying to address.

visual lock-in
prior directions
vision-language models
grounding
representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prior Directions
visual lock-in
vision-language models
representation geometry
controlled intervention