🤖 AI Summary
Modeling the dual spatial and feature-based attention mechanisms observed in human visual cognition remains challenging, particularly in bridging neural network architectures with cognitive principles. Method: We propose a dual-path neural network architecture wherein a primary pathway performs core visual tasks, while an auxiliary pathway dynamically generates top-down attention signals—encoding both spatial localization and feature selectivity (e.g., color, orientation)—and modulates the primary pathway via gated control. Crucially, the model is trained end-to-end without explicit attention supervision. Contribution/Results: For the first time, this approach spontaneously yields cognitively grounded dual attention patterns during training. Interpretability analyses confirm that the learned attention maps exhibit both spatial precision and feature specificity. Consequently, the model demonstrates significantly improved generalization under complex visual scenes, establishing a novel paradigm for computational cognitive modeling grounded in biologically plausible attention mechanisms.
📝 Abstract
Visual attention is a mechanism closely intertwined with vision and memory. Top-down information influences visual processing through attention. We designed a neural network model inspired by aspects of human visual attention. This model consists of two networks: one serves as a basic processor performing a simple task, while the other processes contextual information and guides the first network through attention to adapt to more complex tasks. After training the model and visualizing the learned attention response, we discovered that the model's emergent attention patterns corresponded to spatial and feature-based attention. This similarity between human visual attention and attention in computer vision suggests a promising direction for studying human cognition using neural network models.