multi-scale cross-layer attention

Designs and implements attention modules that compute weighted interactions both across spatial scales and between network layers—covering multi-scale self-attention, cross-layer/layer attention, and windowed or shifted-window variants (e.g., Swin-style). Uses these modules to fuse multi-resolution feature maps, emphasize fine-grained details, and analyze how cross-layer aggregation affects representational quality and robustness to degraded inputs.

multi-scalecross-layerattention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Convolutional Rectangular Attention Module

Mar 13, 2025
HN
Hai-Vy Nguyen
🏛️ Ampere Software Technology | Institut de mathématiques de Toulouse | Institut de Recherche en Informatique de Toulouse | Université Côte d'Azur

This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.

Enhances generalization with rectangular attention regions using 5 parameters.Improves model performance by focusing on discriminative image parts.Introduces a spatial attention module for convolutional networks.

Attention-Guided Multi-scale Interaction Network for Face Super-Resolution

Sep 01, 2024
XW
Xujie Wan
🏛️ Nanjing University of Posts and Telecommunications | Soochow University | Beijing University of Posts and Telecommunications | Southeast University | Nanjing University of Science and Technology | National Tsing Hua University

To address insufficient and weakly complementary multi-scale feature fusion in face super-resolution (FSR), this paper proposes a hybrid CNN-Transformer architecture. Our method introduces three key innovations: (1) a Local-Global Feature Interaction (LGFI) module that enables cross-scale and cross-stage collaborative modeling of local details and global structure; (2) a Selective Kernel Attention Fusion (SKAF) module that adaptively weights and fuses multi-scale features via kernel-wise attention; and (3) Residual Deep Feature Extraction (RDFE), which enhances deep representation learning through hierarchical residual connections. Evaluated on standard benchmarks—including CelebA and FFHQ—our approach achieves state-of-the-art (SOTA) performance, significantly improving high-frequency texture recovery and visual fidelity. Moreover, it reduces computational overhead and accelerates inference speed without compromising accuracy.

Adaptively select feature fusions across different network phasesEnhance complementarity of global and local features in FSRFuse multi-scale features in hybrid networks for face super-resolution

Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task Arithmetic

Jul 09, 2024
RJ
Ruochen Jin
🏛️ East China Normal University | University of Pennsylvania

In task arithmetic, multi-task weight coupling induces interference, degrading both training efficiency and generalization. Method: We propose a novel paradigm that fine-tunes only the attention modules of Transformers—revealing, for the first time, their intrinsic kernel-like behavior. Through systematic analysis, we identify that representation modules facilitate weight decoupling, whereas task-specific heads impede it, thereby establishing a modular decoupling design principle. Contribution/Results: Our method enhances decoupling and zero-shot task generalization without additional training. It significantly outperforms baselines across multiple benchmarks while avoiding the double training overhead required by Neural Tangent Kernel (NTK) linearization. Crucially, it achieves superior weight decoupling and single-task performance, offering a more efficient and effective alternative to existing linearized or fully fine-tuned approaches.

Model InterferenceMulti-task LearningWeighting Strategies

Scratching Visual Transformer's Back with Uniform Attention

Oct 16, 2022
NH
Nam Hyeon-Woo
🏛️ POSTECH | NAVER AI Lab | Tübingen University | Yonsei University

Vision Transformers (ViTs) suffer from implicit preference for dense spatial interactions in their multi-head self-attention (MSA), leading to steep softmax gradients, optimization instability, and limited generalization—contrary to conventional intuition. This work is the first to identify and characterize this phenomenon. We propose Context Broadcasting (CB), a zero-parameter, single-line-deployable mechanism that injects uniform dense attention at each layer to explicitly guide MSA toward satisfying softmax constraints. CB imposes no additional parameters or computational overhead; instead, it reshapes the attention density distribution via lightweight context broadcasting. Evaluated on ImageNet and other benchmarks, CB consistently enhances ViT capacity and generalization, delivering stable accuracy gains across architectures. Our approach offers both a novel analytical lens into Transformer attention dynamics and a practical, implementation-friendly tool for improving attention behavior without architectural modification.

Dense Attention MapsModel AdaptabilityVisual Transformers

Demystify Transformers & Convolutions in Modern Image Deep Networks

Nov 10, 2022
JD
Jifeng Dai
🏛️ Tsinghua University | Shanghai Artificial Intelligence Laboratory | Huazhong University of Science and Technology | Fudan University | The Chinese University of Hong Kong | SenseTime Research | South China University of Technology

This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.

Analyze performance differences in attention vs convolutionCompare spatial token mixers in vision backbonesUnify architecture to isolate feature transformation effects

Latest Papers

What's happening recently
View more

This work investigates the discrepancy between weight ablation and activation patching outcomes in real pretrained models, extending residual block interaction analysis to multi-layer settings. By decomposing interactions via double integrals into intra-block terms and cross-layer remainders, it derives—for the first time—a closed-form upper bound on the Jacobian norm of attention submodules, explicitly characterizing the curvature constant. Combining weight ablation, activation patching, and mixed second-derivative analysis, the study validates that the theoretical bounds hold without violation across all tested cases on Qwen2.5-1.5B-Instruct. Furthermore, it identifies a three-layer indirect object identification circuit shared across five diverse instances, revealing substantial cross-layer interactions that transcend the limitations of single-block theoretical frameworks.

activation patchingattention Jacobiancross-layer interaction

This work addresses the limitations of existing layer attention mechanisms—namely high computational complexity, static information updating, and inadequate modeling of long-range dependencies—by proposing Key-Correlated Layer Attention (KCLA). KCLA leverages the high cosine similarity among inter-layer Key representations to establish a dynamic cross-layer interaction mechanism with linear time complexity and constant space complexity. By integrating key-correlation-based linear attention with adaptive information fusion, KCLA maintains strong long-range dependency modeling capabilities while achieving computational efficiency independent of network depth. Experimental results demonstrate consistent and significant performance improvements across diverse tasks, including image classification, object detection, and medical image segmentation.

computational complexitycross-layer dependencyinter-layer interaction

This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.

efficient attentionfull attentionhybrid architectures

Existing stereo image super-resolution methods struggle to effectively integrate intra-view details with inter-view complementary information. To address this challenge, this work proposes a multi-scale interaction network that enhances single-view feature representation through a multi-scale spatial-channel attention module and introduces a dual-view epipolar attention module to achieve precise cross-view matching along epipolar lines by incorporating epipolar geometric constraints with an optimal transport algorithm. The method innovatively combines large separable convolutional kernel attention with a geometry-guided optimal transport mechanism. Extensive experiments demonstrate that the proposed approach outperforms most state-of-the-art methods on multiple benchmark datasets, and ablation studies confirm the effectiveness of each component.

binocular systemscross-view informationepipolar line

This work addresses the entanglement of routing and filtering functions in conventional attention mechanisms, which leads to structural opacity and optimization challenges. The authors propose S-D Attention, which explicitly decouples the interaction matrix into a low-rank routing component and a symmetric filtering component, thereby clearly distinguishing these two mechanisms for the first time. They further uncover that routing self-organizes into a spectral cascade phenomenon in deep networks. Leveraging this insight, they achieve stable training without layer normalization and validate their approach using linear attention variants (e.g., ELU+1) and effective rank analysis. Experiments show that linearizing the first seven layers of a 125M-parameter model incurs less than 5% perplexity degradation, while cascade-informed architectures reduce attention parameters by 47%–65% with only a 3.9%–8.4% increase in perplexity.

attentionattention decompositionfiltering

Hot Scholars

TT

Truyen Tran

Professor | Head of AI, Health and Science @ Deakin University
artificial intelligenceAI for healthAI for science
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
MB

Mohammed Bennamoun

Winthrop Professor - University of Western Australia
Artificial IntelligenceComputer VisionDeep LearningFace Recognition
YZ

Yidan Zhang

PhD Student, the Chinese University of Hong Kong, Shenzhen
computer visiondeep learning
JY

Jiayu Yang

The Australian National University
3D Computer Vision3D AIGC3D ReconstructionMulti-view Stereo