Score
Design and implement classification models that integrate attention mechanisms (e.g., attention-guided/attention-augmented layers and convolutional block attention modules) to modulate feature representations by emphasizing informative spatial and channel regions and suppressing irrelevant inputs. Build and evaluate architectures and training pipelines that incorporate these attention modules to improve classification accuracy and robustness across datasets.
This study systematically investigates the impact of attention mechanisms on CNN-based image classification performance. We integrate two lightweight attention modules—SE and CBAM—into a ResNet backbone and conduct comparative experiments on CIFAR-10 and an ImageNet subset. For the first time, we quantitatively characterize the accuracy–efficiency trade-off introduced by attention: a 1.8% top-1 accuracy gain on CIFAR-10 incurs a 23% increase in inference latency, with gains scaling significantly with task complexity. We further propose a low-overhead attention embedding scheme that preserves module generality while alleviating computational bottlenecks. Results demonstrate that attention primarily enhances global contextual modeling to compensate for the limited local receptive fields inherent in standard CNNs. Consequently, its effectiveness is highly contingent upon both semantic complexity of the target task and real-time inference constraints.
Existing EEG signal modeling in brain–computer interfaces (BCIs) suffers from insufficient robustness in multidimensional (temporal–spectral–spatial) representation learning and multimodal fusion. Method: This paper systematically analyzes and unifies the embedding paradigms and applicability boundaries of conventional attention and Transformer-based multi-head self-attention for EEG feature modeling. It proposes the first BCI-robust cross-modal attention fusion framework integrating CNNs, RNNs, and Transformers, incorporating three-dimensional attention (channel-, time-, and frequency-wise) alongside multimodal cross-attention mechanisms. Contribution/Results: Experiments demonstrate substantial improvements in EEG representation capacity and generalization: average classification accuracy increases by 5.2–8.7% on motor imagery and emotion recognition benchmarks, while robustness against various noise perturbations is significantly enhanced.
Deep neural networks lack biologically plausible selective attention mechanisms, limiting both efficiency and accuracy in image recognition. To address this, we propose a spatial attention module grounded in information bottleneck theory. Our method explicitly optimizes mutual information: it minimizes the mutual information between the attention representation and the input to suppress redundancy, while maximizing the mutual information between the attention representation and task labels to enhance discriminability. Crucially, we introduce learnable anchors to quantize continuous attention scores—a novel design that strengthens information constraints and improves interpretability of attention maps. By integrating variational attention modeling with deep network embedding, our approach achieves significant performance gains across image classification, fine-grained recognition, and cross-domain classification tasks. The resulting attention maps exhibit high discriminability, strong background suppression, and enhanced interpretability.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
This work addresses the poor conditioning of the Jacobian matrix in Transformer attention mechanisms, which often leads to training instability and performance degradation. For the first time, it explicitly establishes a theoretical link between the condition number of the attention Jacobian and the spectral properties of the query, key, and value projection matrices. Building on this insight, the paper proposes a general, plug-and-play spectral regularization strategy that improves the Jacobian’s condition number by optimizing the singular value distribution of these projection matrices. Notably, the method requires no architectural modifications and consistently enhances performance across diverse Transformer variants and tasks, demonstrating both its effectiveness and broad applicability.
This study investigates the impact of various attention mechanisms on the performance of 3D video classification models under a setting where spatial resolution is enhanced at the expense of temporal cues. Building upon three canonical 3D CNN architectures—MC3, R3D, and R(2+1)D—the authors introduce Dropout layers to simulate scenarios with limited temporal information and systematically integrate ten attention modules, including CBAM, TCN, multi-head attention, and channel attention, for comparative evaluation. Experiments on the UCF101 dataset demonstrate that an enhanced R(2+1)D model combined with multi-head attention achieves 88.98% accuracy, while revealing substantial performance variations across different attention mechanisms at the class level. This work provides the first systematic analysis of how degraded temporal features critically affect 3D action recognition, offering novel insights for modeling high-resolution videos with weak temporal signals.
This study addresses the challenge of inefficient multimodal (visual-auditory) module coordination in assistive perception systems. We propose a lightweight, domain-specific modular deep learning framework: a CNN processes eye-region images for gaze-state estimation; a deeper CNN models facial expressions (trained on FER2013); and a CNN-LSTM hybrid architecture performs speaker identification (using a custom audio dataset). Each module is independently optimized and designed for plug-and-play integration. Our key contribution is empirically validating that high-accuracy unimodal modeling combined with a loosely coupled modular architecture achieves superior performance and deployment flexibility under resource constraints. Experiments yield accuracies of 93.0% (gaze state), 97.8% (facial expression), and 96.89% (speaker ID), significantly outperforming end-to-end joint modeling baselines. This work establishes a scalable, modular paradigm for assistive technologies.
This work addresses the lack of systematic analysis and unified design principles in existing channel-spatial attention fusion strategies. Under a consistent experimental framework, the authors construct and comprehensively evaluate 18 channel-spatial attention topologies, spanning serial, parallel, multi-scale, and residual architectures. Extensive experiments across diverse vision and medical imaging datasets reveal a coupling relationship among data scale, architectural design, and performance. The study proposes practical guidelines for attention module construction tailored to data regime size: cascaded channel–multi-scale spatial attention excels in small-sample tasks; learnable parallel fusion achieves optimal results at medium scales; and large-scale scenarios benefit from parallel structures augmented with dynamic gating. Additionally, the work validates the advantage of spatial-before-channel ordering for fine-grained classification and demonstrates the efficacy of residual connections in mitigating gradient vanishing.
Manual verification in wide-field solar system surveys severely limits the efficiency of moving object detection. Method: This paper proposes an end-to-end multi-input convolutional neural network (CNN) architecture incorporating the Convolutional Block Attention Module (CBAM), enabling direct processing of multi-depth, multi-frame stacked images and adaptive enhancement of salient motion features in both spatial and channel dimensions. The method eliminates traditional manual screening, achieving fully automated detection and classification from raw image sequences. Results: Evaluated on approximately 2,000 real survey images, the model achieves 98.9% accuracy and an AUC of 0.992. With optimized detection thresholds, it maintains high recall while reducing human verification effort by over 99%. This work significantly advances the automation level and operational efficiency of moving object discovery in astronomical surveys.
This work addresses the limitations of existing global workspace architectures, which lack effective attention mechanisms and struggle to balance noise robustness with cross-task generalization in multimodal fusion. Inspired by cognitive neuroscience, we propose the first explicit modality selection attention mechanism tailored for global workspace frameworks. Our approach employs a learnable, top-down attention process to dynamically integrate and select relevant modality-specific information. Evaluated on the Simple Shapes and MM-IMDb 1.0 datasets, the proposed method significantly enhances robustness to input noise while achieving performance on par with state-of-the-art approaches on MM-IMDb 1.0. These results demonstrate its superior capacity for cross-task and cross-modal generalization, highlighting the efficacy of biologically inspired attention in multimodal reasoning systems.