GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding

📅 2025-11-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing gaze understanding methods suffer from task fragmentation and insufficient cross-modal collaboration. To address this, we propose the first unified vision-language multitask framework that jointly performs person detection, gaze target localization, and salient object recognition, with selective subtask execution enabled by text prompts. Our approach innovatively integrates vision-language models into gaze understanding, introduces an object-level gaze detection metric $AP_{ob}$, and fuses RGB and HHA depth features via a Transformer for cross-modal alignment and joint modeling. Extensive experiments on GazeFollow and VideoAttentionTarget demonstrate state-of-the-art performance, achieving significant improvements in multitask collaborative reasoning and flexible task scheduling.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Gaze understanding unifies the detection of people, their gaze targets, and objects of interest into a single framework, offering critical insight into visual attention and intent estimation. Although prior research has modelled gaze cues in visual scenes, a unified system is still needed for gaze understanding using both visual and language prompts. This paper introduces GazeVLM, a novel Vision-Language Model (VLM) for multi-task gaze understanding in images, addressing person detection, gaze target detection, and gaze object identification. While other transformer-based methods exist for gaze analysis, GazeVLM represents, to our knowledge, the first application of a VLM to these combined tasks, allowing for selective execution of each task. Through the integration of visual (RGB and depth) and textual modalities, our ablation study on visual input combinations revealed that a fusion of RGB images with HHA-encoded depth maps, guided by text prompts, yields superior performance. We also introduce an object-level gaze detection metric for gaze object identification ($AP_{ob}$). Through experiments, GazeVLM demonstrates significant improvements, notably achieving state-of-the-art evaluation scores on GazeFollow and VideoAttentionTarget datasets.
Problem

Research questions and friction points this paper is trying to address.

Unified framework for detecting people and gaze targets
First vision-language model for multi-task gaze understanding
Integrates visual and textual cues for intent estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified vision-language model for multi-task gaze understanding
Fusion of RGB and HHA depth maps with text prompts
Introduces object-level gaze detection metric AP_ob
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Elm Company
A
Athul M. Mathew
Applied Research, Elm Company, Saudi Arabia
H
Haithem Hermassi
Applied Research, Elm Company, Saudi Arabia
Thariq Khalid
Thariq Khalid
Elm Company
Deep LearningComputer VisionNLPArtificial Intelligence
A
A. Khan
Applied Research, Elm Company, Saudi Arabia
R
R. Souissi
Applied Research, Elm Company, Saudi Arabia