🤖 AI Summary
This work addresses the limitation of existing visual saliency models, which predominantly rely on the free-viewing assumption and thus struggle to capture task-driven attentional patterns. To overcome this, the authors propose a task-driven saliency prediction model that explicitly integrates natural language descriptions of task semantics into the visual attention mechanism for the first time, establishing a task-conditioned architecture for saliency prediction. Experimental results demonstrate that the proposed approach effectively captures the dynamic shifts in human attention across different tasks and significantly improves prediction accuracy in task-oriented viewing scenarios compared to conventional methods.
📝 Abstract
Visual saliency aims to predict the regions of an image most likely to attract human visual attention. While most saliency models assume free-viewing conditions, human attention is often shaped by explicit task goals. In this work, we address task-driven saliency prediction by proposing a model that conditions visual attention on natural-language task descriptions. The model produces task-dependent saliency maps that reflect how attention shifts under different viewing intents. Through quantitative and qualitative analysis, we show that incorporating explicit task semantics enables more faithful modeling of goal-directed visual attention.