🤖 AI Summary
This work addresses the ambiguity in large vision-language models (VLMs) caused by vague user queries in open-domain visual question answering (VQA). To resolve this, the authors propose a training-free, inference-time method that leverages real-time eye-tracking data—specifically, gaze fixations captured at the onset of user questioning—to guide VLMs toward relevant image regions and thereby infer user intent more accurately. The study introduces the first VQA interaction protocol and benchmark that integrates eye-movement data. Evaluated on 500 image-question pairs involving ambiguous queries, the approach significantly improves answer accuracy from 35.2% to 77.2%, with consistent performance gains observed across multiple mainstream VLM architectures.
📝 Abstract
We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ambiguity in open-ended VQA. Through a comprehensive user study with 500 unique image-question pairs, we demonstrate that fixations closest to the time participants start verbally asking their questions are the most informative for disambiguation in Large VLMs, more than doubling the accuracy of responses on ambiguous questions (from 35.2% to 77.2%) while maintaining performance on unambiguous queries. We evaluate our approach across state-of-the-art VLMs, showing consistent improvements when gaze data is incorporated in ambiguous image-question pairs, regardless of architectural differences. We release a new benchmark dataset to use eye movement data for disambiguated VQA, a novel real-time interactive protocol, and an evaluation suite.