IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models

📅 2026-02-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the ambiguity in large vision-language models (VLMs) caused by vague user queries in open-domain visual question answering (VQA). To resolve this, the authors propose a training-free, inference-time method that leverages real-time eye-tracking data—specifically, gaze fixations captured at the onset of user questioning—to guide VLMs toward relevant image regions and thereby infer user intent more accurately. The study introduces the first VQA interaction protocol and benchmark that integrates eye-movement data. Evaluated on 500 image-question pairs involving ambiguous queries, the approach significantly improves answer accuracy from 35.2% to 77.2%, with consistent performance gains observed across multiple mainstream VLM architectures.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Question Answering

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ambiguity in open-ended VQA. Through a comprehensive user study with 500 unique image-question pairs, we demonstrate that fixations closest to the time participants start verbally asking their questions are the most informative for disambiguation in Large VLMs, more than doubling the accuracy of responses on ambiguous questions (from 35.2% to 77.2%) while maintaining performance on unambiguous queries. We evaluate our approach across state-of-the-art VLMs, showing consistent improvements when gaze data is incorporated in ambiguous image-question pairs, regardless of architectural differences. We release a new benchmark dataset to use eye movement data for disambiguated VQA, a novel real-time interactive protocol, and an evaluation suite.
Problem

Research questions and friction points this paper is trying to address.

visual question answering
ambiguity resolution
large vision-language models
eye-tracking
open-ended VQA
Innovation

Methods, ideas, or system contributions that make the work stand out.

inference-time saccades
gaze-guided disambiguation
training-free VQA
eye-tracking for VLMs
open-ended visual question answering
🔎 Similar Papers
P
Parsa Madinei
Department of Computer Science, UC Santa Barbara; Department of Psychological & Brain Sciences, UC Santa Barbara
S
Srijita Karmakar
Department of Psychological & Brain Sciences, UC Santa Barbara
Russell Cohen Hoffing
Russell Cohen Hoffing
US DEVCOM ARL
F
Felix Gervitz
DEVCOM Army Research Laboratory
M
Miguel P. Eckstein
Department of Computer Science, UC Santa Barbara; Department of Psychological & Brain Sciences, UC Santa Barbara