Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection

📅 2025-04-27
🏛️ Gaze Meets ML
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the gaze target detection (GTD) task in natural images. We propose a depth-enhanced, multimodal end-to-end prediction framework. First, monocular depth estimation is employed to construct 3D representations of both the human subject and the scene. Second, a depth-injected saliency module is designed to enable semantic-aware, gaze-subject-oriented saliency modeling. Third, a multimodal adaptive fusion mechanism jointly encodes facial orientation, geometric depth, and visual saliency features. Our method achieves new state-of-the-art performance on three major benchmarks—VideoAttentionTarget, GazeFollow, and GOO-Real. Ablation studies confirm the effectiveness and complementarity of each component. The core contributions are twofold: (1) the first explicit integration of depth cues into saliency modeling for GTD, and (2) the establishment of a multimodal collaborative learning paradigm specifically tailored to GTD.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
Gaze target detection (GTD) is the task of predicting where a person in an image is looking. This is a challenging task, as it requires the ability to understand the relationship between the person's head, body, and eyes, as well as the surrounding environment. In this paper, we propose a novel method for GTD that fuses multiple pieces of information extracted from an image. First, we project the 2D image into a 3D representation using monocular depth estimation. We then extract a depth-infused saliency module map, which highlights the most salient ( extit{attention-grabbing}) regions in image for the subject in consideration. We also extract face and depth modalities from the image, and finally fuse all the extracted modalities to identify the gaze target. We quantitatively evaluated our method, including the ablation analysis on three publicly available datasets, namely VideoAttentionTarget, GazeFollow and GOO-Real, and showed that it outperforms other state-of-the-art methods. This suggests that our method is a promising new approach for GTD.
Problem

Research questions and friction points this paper is trying to address.

Predicting gaze targets in images using multi-modal fusion
Integrating depth, saliency, and face modalities for accurate detection
Outperforming state-of-the-art methods on three benchmark datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fuses multi-modal data for gaze detection
Uses depth-infused saliency module maps
Projects 2D images into 3D representations
🔎 Similar Papers
No similar papers found.
Elm Company
A
Athul M. Mathew
Elm Company, Saudi Arabia
Arshad Khan
Arshad Khan
Elm Company, Saudi Arabia
Thariq Khalid
Thariq Khalid
Elm Company
Deep LearningComputer VisionNLPArtificial Intelligence
F
F. Al-Tam
Elm Company, Saudi Arabia
R
R. Souissi
Elm Company, Saudi Arabia