radiomics attention gate

Designs, implements, or analyzes attention-gating modules that fuse radiomic descriptors (e.g., GLCM, LBP and other texture features) with convolutional feature maps to modulate skip-connection attention maps and emphasize texture-driven regions; the component is constructed to both improve texture-sensitive segmentation/decision signals and provide ante-hoc interpretability cues.

radiomicsattentiongate

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional CNNs struggle to model fine-grained and complex discriminative features in medical image analysis. To address this, we propose a systematic attention module integration framework that unifies Squeeze-and-Excitation and hybrid convolutional attention mechanisms into five mainstream architectures—VGG16, ResNet18, InceptionV3, DenseNet121, and EfficientNetB5—enabling adaptive feature recalibration along both channel and spatial dimensions. The method is validated across modalities—brain tumor MRI and histopathological images—demonstrating significant improvements in classification accuracy and localization interpretability; EfficientNetB5 augmented with hybrid attention achieves state-of-the-art performance. Our core contributions are: (1) establishing a reproducible, plug-and-play attention evaluation paradigm; and (2) empirically validating the generalizable performance gains and enhanced mechanistic interpretability of attention mechanisms across diverse architectures and medical imaging modalities.

Enhancing generalization across heterogeneous medical imaging modalitiesImproving feature focus and discriminative performance in CNNsIntegrating attention modules into CNNs for enhanced medical image diagnosis

Convolutional Rectangular Attention Module

Mar 13, 2025
HN
Hai-Vy Nguyen
🏛️ Ampere Software Technology | Institut de mathématiques de Toulouse | Institut de Recherche en Informatique de Toulouse | Université Côte d'Azur

This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.

Enhances generalization with rectangular attention regions using 5 parameters.Improves model performance by focusing on discriminative image parts.Introduces a spatial attention module for convolutional networks.

This study addresses the limitations of deep learning in medical image segmentation—namely, poor interpretability, excessive parameter count, and insufficient clinical trustworthiness—by proposing RadiomicNet, a lightweight dual-stream architecture that integrates handcrafted radiomic features. The method introduces a novel Radiomic Attention Gate (RAG) to inject Gray-Level Co-occurrence Matrix (GLCM) and Local Binary Pattern (LBP) features into the skip connections of a MobileNetV2 encoder-decoder framework, alongside a radiomic consistency loss to improve prediction calibration. With only 3.27 million parameters, RadiomicNet achieves Dice scores of 0.763 and 0.854 on the BUSI and Kvasir-SEG datasets, respectively, significantly outperforming U-KAN while providing inherent interpretability and explicit quantification of key feature contributions.

deep learninginterpretabilitymedical image segmentation

Demystify Transformers & Convolutions in Modern Image Deep Networks

Nov 10, 2022
JD
Jifeng Dai
🏛️ Tsinghua University | Shanghai Artificial Intelligence Laboratory | Huazhong University of Science and Technology | Fudan University | The Chinese University of Hong Kong | SenseTime Research | South China University of Technology

This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.

Analyze performance differences in attention vs convolutionCompare spatial token mixers in vision backbonesUnify architecture to isolate feature transformation effects

Information Bottleneck Approach to Spatial Attention Learning

Aug 01, 2021
QL
Qiuxia Lai
🏛️ The Chinese University of Hong Kong | University of Electronic Science and Technology of China

Deep neural networks lack biologically plausible selective attention mechanisms, limiting both efficiency and accuracy in image recognition. To address this, we propose a spatial attention module grounded in information bottleneck theory. Our method explicitly optimizes mutual information: it minimizes the mutual information between the attention representation and the input to suppress redundancy, while maximizing the mutual information between the attention representation and task labels to enhance discriminability. Crucially, we introduce learnable anchors to quantize continuous attention scores—a novel design that strengthens information constraints and improves interpretability of attention maps. By integrating variational attention modeling with deep network embedding, our approach achieves significant performance gains across image classification, fine-grained recognition, and cross-domain classification tasks. The resulting attention maps exhibit high discriminability, strong background suppression, and enhanced interpretability.

Deep Neural NetworksImage RecognitionSelective Attention

Latest Papers

What's happening recently
View more

This study addresses the challenge of accurately predicting early human fixation regions in visual search tasks with unknown target locations, aiming to model bottom-up visual attention allocation. To this end, it proposes two multi-feature fusion pipelines that systematically integrate structure-oriented Gabor filter responses with statistical texture features derived from the Gray-Level Co-occurrence Matrix (GLCM)—a novel combination in this context. The approach is validated on digital breast tomosynthesis images, demonstrating that the generated salient regions exhibit strong alignment with human observers’ early eye movements and outperform conventional threshold-based models. These findings highlight the complementary roles of Gabor and GLCM features in visual information encoding and offer a new pathway for developing perception-driven observer models.

fixation predictionGabor featuresregion of interest

This study investigates whether Gram matrix–based texture representations in convolutional neural networks (CNNs) align with human texture perception and whether such alignment improves as CNNs become better models of the visual system. By integrating human psychophysical data with Brain-Score neural benchmarking, the authors conduct a cross-model comparative analysis across diverse CNN architectures. The findings reveal that, regardless of a model’s performance in object recognition or neural predictivity, its texture representations consistently fail to accurately capture key aspects of human perceptual judgments. This work demonstrates, for the first time, no significant association between a CNN’s general visual modeling capability and its fidelity in representing human texture perception, thereby challenging the prevailing assumption that object recognition training inherently yields human-like texture representations. The results suggest that human texture perception likely relies on contextual integration mechanisms beyond local feature correlations encoded by standard CNNs.

convolutional neural networksGram matricesperceptual alignment

This study investigates the alignment between deep visual models and the human visual system in texture perception, moving beyond the conventional object-centric paradigm. By constructing stimulus sets of varying complexity using three source-image-based texture synthesis algorithms and leveraging human psychophysical data, the authors systematically compare the internal representational properties of convolutional neural networks (CNNs) and Vision Transformers (ViTs). The findings reveal that ViTs exhibit highly consistent texture representations that are insensitive to stimulus complexity and significantly better predict human texture discrimination behavior than CNNs. These results underscore network architecture as a critical determinant of texture representation, suggesting that ViTs more closely approximate human visual mechanisms in processing texture information.

deep vision modelshuman perceptionrepresentation alignment

This work addresses the limited spatial accuracy and causal fidelity of attention mechanisms in existing vision models, which often fail to align with true discriminative regions. To this end, the authors propose CAMAL, a novel approach that leverages segmentation masks as supervision signals to regularize attention during training, explicitly guiding it toward relevant regions while suppressing irrelevant ones. Built upon class activation maps, CAMAL integrates segmentation masks into a joint deep reinforcement learning framework, enhancing both spatial alignment and faithfulness of attention without incurring additional inference overhead. Experimental results demonstrate that CAMAL consistently improves alignment performance across various settings, achieving over a 35% gain in attention faithfulness while maintaining or even improving model generalization.

attention alignmentattention faithfulnessmodel explainability

This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.

attention mechanismmodel interpretabilitysaliency maps

Hot Scholars

UO

Utku Ozbulak

Research Professor at Ghent University
Trustworthy AIMedical imagingBiomedical imagingSelf-supervised learning