OpenVAM: Open-World Visual Attention Modeling with VLMs

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited interpretability and constrained cross-domain robustness in visual attention prediction by proposing a unified framework for multi-domain saliency localization and semantic interpretation. Methodologically, we design a decoupled alignment architecture that leverages a dense visual pathway and a linguistic semantic head to achieve precise localization and attribution. Furthermore, we construct a three-stage training strategy incorporating parameter-efficient fine-tuning alongside a large-scale, multi-domain annotation generation pipeline. This work significantly enhances the model's cross-domain generalization capability and achieves, for the first time, image-grounded interpretable attention prediction.
📝 Abstract
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web layouts. We, therefore, introduce OpenVAM (Open-world Visual Attention Modeling with VLMs), a unified framework that jointly addresses universality and explainability across heterogeneous domains (natural scenes, commercial imagery, and UI/web layouts) and supervision modalities. OpenVAM adopts a decoupled-but-aligned design: a dedicated dense visual pathway provides stable, spatially precise localization, while an instruction-following vision--language semantic head generates grounded what/why explanations conditioned on the same image and data-type prompts. A three-stage training strategy preserves strong localization priors while progressively introducing language grounding and improving explanation alignment via parameter-efficient adaptation without perturbing the saliency branch. We further propose a scalable pipeline to generate multi-domain saliency-reason annotations for training and systematic evaluation. Experiments across diverse datasets show that OpenVAM improves robustness under domain shift while producing image-grounded explanations that make saliency predictions more interpretable.
Problem

Research questions and friction points this paper is trying to address.

visual attention modeling
saliency prediction
explainability
domain shift
open-world
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Attention Modeling
Vision-Language Models
Explainable Saliency
Parameter-Efficient Adaptation
Cross-Domain Robustness
🔎 Similar Papers
No similar papers found.