attention-based feature refinement

Designs and implements attention-based modules and upsampling operators that refine spatial feature maps in neural networks, using attention gates or attention-guided upsampling to amplify weak local cues, suppress irrelevant background activations, and stabilize representations under appearance changes. Analyzes and validates these components’ effects on decoder-stage feature refinement and final dense predictions, including their ability to preserve global structure and focus on salient spatial regions.

attention-basedfeaturerefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing cross-attention-based methods for generating high-resolution, pixel-dense features incur substantial computational overhead. To address this limitation, this work proposes UPLiFT, a task-agnostic, lightweight architecture that combines iterative upsampling with a local attention mechanism. The core innovation lies in the Local Attender operator, which employs fully local attention pooling to effectively circumvent the traditional trade-off between stability and efficiency in iterative upsampling. UPLiFT achieves state-of-the-art performance across multiple dense prediction tasks while significantly reducing inference cost compared to existing approaches. Notably, in VAE feature upsampling, it matches the performance of cutting-edge models such as Coupled Flow Matching, demonstrating both efficacy and efficiency.

efficient inferencefeature upsamplingpixel-dense features

Convolutional Rectangular Attention Module

Mar 13, 2025
HN
Hai-Vy Nguyen
🏛️ Ampere Software Technology | Institut de mathématiques de Toulouse | Institut de Recherche en Informatique de Toulouse | Université Côte d'Azur

This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.

Enhances generalization with rectangular attention regions using 5 parameters.Improves model performance by focusing on discriminative image parts.Introduces a spatial attention module for convolutional networks.

A Neural Network Model of Spatial and Feature-Based Attention

Jun 05, 2025
RH
Ruoyang Hu
🏛️ University of Rochester

Modeling the dual spatial and feature-based attention mechanisms observed in human visual cognition remains challenging, particularly in bridging neural network architectures with cognitive principles. Method: We propose a dual-path neural network architecture wherein a primary pathway performs core visual tasks, while an auxiliary pathway dynamically generates top-down attention signals—encoding both spatial localization and feature selectivity (e.g., color, orientation)—and modulates the primary pathway via gated control. Crucially, the model is trained end-to-end without explicit attention supervision. Contribution/Results: For the first time, this approach spontaneously yields cognitively grounded dual attention patterns during training. Interpretability analyses confirm that the learned attention maps exhibit both spatial precision and feature specificity. Consequently, the model demonstrates significantly improved generalization under complex visual scenes, establishing a novel paradigm for computational cognitive modeling grounded in biologically plausible attention mechanisms.

Bridging human cognition and neural networksIntegrating spatial and feature-based attentionModeling human visual attention mechanisms

This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.

Attention MechanismLarge Vision-Language ModelsModel Redundancy

Demystify Transformers & Convolutions in Modern Image Deep Networks

Nov 10, 2022
JD
Jifeng Dai
🏛️ Tsinghua University | Shanghai Artificial Intelligence Laboratory | Huazhong University of Science and Technology | Fudan University | The Chinese University of Hong Kong | SenseTime Research | South China University of Technology

This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.

Analyze performance differences in attention vs convolutionCompare spatial token mixers in vision backbonesUnify architecture to isolate feature transformation effects

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic analysis and unified design principles in existing channel-spatial attention fusion strategies. Under a consistent experimental framework, the authors construct and comprehensively evaluate 18 channel-spatial attention topologies, spanning serial, parallel, multi-scale, and residual architectures. Extensive experiments across diverse vision and medical imaging datasets reveal a coupling relationship among data scale, architectural design, and performance. The study proposes practical guidelines for attention module construction tailored to data regime size: cascaded channel–multi-scale spatial attention excels in small-sample tasks; learnable parallel fusion achieves optimal results at medium scales; and large-scale scenarios benefit from parallel structures augmented with dynamic gating. Additionally, the work validates the advantage of spatial-before-channel ordering for fine-grained classification and demonstrates the efficacy of residual connections in mitigating gradient vanishing.

Attention FusionChannel AttentionParallel Design

This work addresses the attention misalignment and error accumulation in foundational segmentation models during closed-loop iterative prompting, which stem from decoder coupling drift. The study is the first to formally characterize this phenomenon by modeling iterative prompting as a discrete-time dynamical system. Building upon the SAM architecture, the authors propose a training-free inference-time stabilization framework that leverages truth-agnostic prompt–image coupling metrics, attention stability analysis, and a proximal anchoring strategy to constrain prompt updates across iterations, thereby preserving decoder coupling consistency. Experiments on volumetric electron microscopy data demonstrate that the proposed method significantly enhances attention stability, temporal consistency, and segmentation accuracy, outperforming existing iterative prompting approaches.

attention alignmentclosed-loop segmentationdecoder coupling drift

Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.

annotation structureobject detectionrepresentation robustness

This work investigates how to explicitly control and quantify the information flow within the attention mechanism of Vision Transformers to uncover the evolutionary process from independent local patch processing to the emergence of global representations. The authors introduce a variational information bottleneck along all pathways where attention writes into the residual stream, without altering the model architecture. This approach models the transmitted information as an explicit, tunable variable, enabling precise constraint and intervention on internal communication. Experiments on ImageNet-100 demonstrate that the method effectively modulates the degree of attention coordination and reveals a clear relationship between classification performance and information routing strategies, offering a novel perspective for understanding the internal mechanisms of Vision Transformers.

attention mechanisminformation bottleneckinformation flow

This work addresses the limitations of Vision Transformers (ViTs) in dense prediction tasks, where low-resolution feature maps hinder performance and conventional upsampling methods often introduce artifacts such as feature leakage, fragmentation, and blurriness. The authors propose an implicit feature upsampling framework that operates without external image guidance, leveraging intermediate hidden states from the ViT to construct inter-layer continuous queries. This enables high-fidelity, task-agnostic feature prediction at arbitrary spatial coordinates while preserving feature-space alignment and supporting continuous-coordinate inference. Experimental results demonstrate significant improvements over existing image-guided upsampling approaches, with a +3.36 mIoU gain on Cityscapes semantic segmentation and a +8.09 increase in PCK@0.10 on the SPair-71k correspondence benchmark.

dense predictiondepth estimationfeature upsampling

Hot Scholars

MH

Ming-Hsuan Yang

University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence
NS

Nicu Sebe

University of Trento
computer visionmultimedia
FL

Fang-Lue Zhang

Senior Lecturer (Associate Professor), Victoria University of Wellington, New Zealand
Computer GraphicsImage and Video ProcessingVR AR MRComputer Vision
NM

Ngai-Man Cheung

Associate Professor, Singapore University of Technology and Design
Image and signal processingComputer VisionAI
HL

Haodong Li

UC San Diego. Prev: HKUST, ZJU, Tencent.
3DVGenerative ModelsAgents