Score
Designs, implements, and evaluates detection models and components that integrate attention mechanisms—including attention modules, attention-guided training/inference, attention-augmented detectors, and learned “where-to-look” policies—to focus computation on salient image regions. Builds and analyzes the impact of these mechanisms on localization and classification precision (including small or occluded targets), false positive rates, and runtime trade-offs to preserve or optimize real-time inference performance.
Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.
Attention mechanisms in diffusion models lack systematic analysis regarding their functional roles, design principles, and cross-modal/task generalizability. Method: We propose the first unified taxonomy for attention modifications in diffusion models, categorizing improvements by architectural component—e.g., U-Net backbone, cross-attention, spatial/channel-wise attention—and integrating multi-dimensional analysis: architectural characterization, modality-aware comparison, performance attribution, and limitation diagnosis. Contribution/Results: Our analysis reveals distinct contribution pathways of attention to generation quality, training stability, sampling efficiency, and controllable editing. We identify critical bottlenecks: poor scalability, high computational redundancy, and weak theoretical interpretability. The framework provides a structured design guide for attention-augmented diffusion models and motivates future directions—including attention sparsification, modular co-optimization, and interpretable attention modeling—to advance both efficacy and understanding.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
To address the excessive computational overhead introduced by attention mechanisms in real-time object detection—compromising the balance between speed and accuracy—this paper proposes YOLOv12. Methodologically, it introduces (1) a lightweight Area Attention module for efficient local-global self-attention modeling; (2) a Residual Efficient Layer Aggregation Network (RELAN) integrated with FlashAttention to enhance feature reuse and long-range dependency capture; and (3) co-optimized backbone and detection head design. On COCO, YOLOv12 achieves 54.1% mAP—3.2% higher than YOLOv8/v10—with 128 FPS inference speed on an NVIDIA Tesla A100 and 19% lower computational cost. It outperforms state-of-the-art YOLO and DETR-based models, establishing a new latency–accuracy trade-off paradigm for real-time detection.
This study systematically investigates the impact of attention mechanisms on CNN-based image classification performance. We integrate two lightweight attention modules—SE and CBAM—into a ResNet backbone and conduct comparative experiments on CIFAR-10 and an ImageNet subset. For the first time, we quantitatively characterize the accuracy–efficiency trade-off introduced by attention: a 1.8% top-1 accuracy gain on CIFAR-10 incurs a 23% increase in inference latency, with gains scaling significantly with task complexity. We further propose a low-overhead attention embedding scheme that preserves module generality while alleviating computational bottlenecks. Results demonstrate that attention primarily enhances global contextual modeling to compensate for the limited local receptive fields inherent in standard CNNs. Consequently, its effectiveness is highly contingent upon both semantic complexity of the target task and real-time inference constraints.
Existing approaches to modeling human attention are highly fragmented across modalities, scenarios, and tasks, lacking a unified framework. This work proposes AAM, a foundational model for human attention that leverages language prompts and hierarchical embeddings in hyperbolic space to represent attention as a cognitive entailment from general to specific. Furthermore, it introduces a fluid dynamics framework grounded in the Fokker–Planck equation to jointly model attention in both static images and dynamic videos. AAM is the first method to achieve unified attention modeling across image, video, and audio-visual tasks, outperforming state-of-the-art approaches by an average of 6% across 16 benchmarks while accelerating video inference by approximately fourfold.
To address the neglect of human visual attention in first-person video object detection, this paper proposes a gaze-guided Vision Transformer framework. It models eye-tracking trajectories as dynamic attention priors and injects them into the multi-head self-attention mechanism; a gaze-aware attention head importance metric is further introduced to explicitly modulate each head’s response strength to fixated regions. Additionally, depth information is fused to enhance spatial perception. The method achieves significant improvements over non-gaze-aware baselines—up to +3.2–5.7% mAP—on a custom simulator dataset and public benchmarks including Ego4D Ego-Motion and Ego-CH-Gaze. Ablation studies validate the effectiveness of both the gaze-guidance mechanism and the depth-fusion module. This work is the first to systematically demonstrate the interpretability value of eye movement signals for evaluating first-person object detection, establishing a novel paradigm for attention modeling in embodied intelligence.
Existing attention-based sparse image matching models exhibit significant performance variations across different local features, yet the individual contributions of detectors and descriptors remain unclear. This work systematically investigates this issue and reveals that the choice of detector has a far greater impact on matching performance than that of the descriptor. Building on this insight, we propose a general, detector-agnostic zero-shot matching strategy: by fine-tuning a Transformer-based matcher with keypoints aggregated from multiple pre-existing detectors, our approach eliminates the need for retraining on any specific detector. Experiments demonstrate that, in zero-shot settings, our method achieves matching accuracy on novel detectors that matches or even surpasses that of models explicitly trained for those detectors, confirming its effectiveness and strong generalization capability.
This study addresses the challenge of accurately predicting early human fixation regions in visual search tasks with unknown target locations, aiming to model bottom-up visual attention allocation. To this end, it proposes two multi-feature fusion pipelines that systematically integrate structure-oriented Gabor filter responses with statistical texture features derived from the Gray-Level Co-occurrence Matrix (GLCM)—a novel combination in this context. The approach is validated on digital breast tomosynthesis images, demonstrating that the generated salient regions exhibit strong alignment with human observers’ early eye movements and outperform conventional threshold-based models. These findings highlight the complementary roles of Gabor and GLCM features in visual information encoding and offer a new pathway for developing perception-driven observer models.
This study addresses the unclear internal mechanisms underlying object localization in vision-language models (VLMs), which hinders their interpretability and performance improvement. It reveals, for the first time at the layer and attention head granularity, that object localization in LLaVA-1.5 and InternVL-3.5 relies on narrow computational pathways formed by a small subset of specialized attention heads—rather than internal semantic rearrangements—exhibiting a “containerized” mechanism. Through token ablation, attention knockout, and causal mediation analysis, the work demonstrates that localization and classification tasks share early visual processing but are driven by distinct sets of attention heads: in LLaVA, critical heads concentrate in early-to-mid layers, whereas in InternVL, they are distributed across mid-to-late layers.