Score
Designs and implements attention modules that compute weighted interactions both across spatial scales and between network layers—covering multi-scale self-attention, cross-layer/layer attention, and windowed or shifted-window variants (e.g., Swin-style). Uses these modules to fuse multi-resolution feature maps, emphasize fine-grained details, and analyze how cross-layer aggregation affects representational quality and robustness to degraded inputs.
This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.
To address insufficient and weakly complementary multi-scale feature fusion in face super-resolution (FSR), this paper proposes a hybrid CNN-Transformer architecture. Our method introduces three key innovations: (1) a Local-Global Feature Interaction (LGFI) module that enables cross-scale and cross-stage collaborative modeling of local details and global structure; (2) a Selective Kernel Attention Fusion (SKAF) module that adaptively weights and fuses multi-scale features via kernel-wise attention; and (3) Residual Deep Feature Extraction (RDFE), which enhances deep representation learning through hierarchical residual connections. Evaluated on standard benchmarks—including CelebA and FFHQ—our approach achieves state-of-the-art (SOTA) performance, significantly improving high-frequency texture recovery and visual fidelity. Moreover, it reduces computational overhead and accelerates inference speed without compromising accuracy.
In task arithmetic, multi-task weight coupling induces interference, degrading both training efficiency and generalization. Method: We propose a novel paradigm that fine-tunes only the attention modules of Transformers—revealing, for the first time, their intrinsic kernel-like behavior. Through systematic analysis, we identify that representation modules facilitate weight decoupling, whereas task-specific heads impede it, thereby establishing a modular decoupling design principle. Contribution/Results: Our method enhances decoupling and zero-shot task generalization without additional training. It significantly outperforms baselines across multiple benchmarks while avoiding the double training overhead required by Neural Tangent Kernel (NTK) linearization. Crucially, it achieves superior weight decoupling and single-task performance, offering a more efficient and effective alternative to existing linearized or fully fine-tuned approaches.
Vision Transformers (ViTs) suffer from implicit preference for dense spatial interactions in their multi-head self-attention (MSA), leading to steep softmax gradients, optimization instability, and limited generalization—contrary to conventional intuition. This work is the first to identify and characterize this phenomenon. We propose Context Broadcasting (CB), a zero-parameter, single-line-deployable mechanism that injects uniform dense attention at each layer to explicitly guide MSA toward satisfying softmax constraints. CB imposes no additional parameters or computational overhead; instead, it reshapes the attention density distribution via lightweight context broadcasting. Evaluated on ImageNet and other benchmarks, CB consistently enhances ViT capacity and generalization, delivering stable accuracy gains across architectures. Our approach offers both a novel analytical lens into Transformer attention dynamics and a practical, implementation-friendly tool for improving attention behavior without architectural modification.
This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.
This work investigates the discrepancy between weight ablation and activation patching outcomes in real pretrained models, extending residual block interaction analysis to multi-layer settings. By decomposing interactions via double integrals into intra-block terms and cross-layer remainders, it derives—for the first time—a closed-form upper bound on the Jacobian norm of attention submodules, explicitly characterizing the curvature constant. Combining weight ablation, activation patching, and mixed second-derivative analysis, the study validates that the theoretical bounds hold without violation across all tested cases on Qwen2.5-1.5B-Instruct. Furthermore, it identifies a three-layer indirect object identification circuit shared across five diverse instances, revealing substantial cross-layer interactions that transcend the limitations of single-block theoretical frameworks.
This work addresses the limitations of existing layer attention mechanisms—namely high computational complexity, static information updating, and inadequate modeling of long-range dependencies—by proposing Key-Correlated Layer Attention (KCLA). KCLA leverages the high cosine similarity among inter-layer Key representations to establish a dynamic cross-layer interaction mechanism with linear time complexity and constant space complexity. By integrating key-correlation-based linear attention with adaptive information fusion, KCLA maintains strong long-range dependency modeling capabilities while achieving computational efficiency independent of network depth. Experimental results demonstrate consistent and significant performance improvements across diverse tasks, including image classification, object detection, and medical image segmentation.
This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.
Existing stereo image super-resolution methods struggle to effectively integrate intra-view details with inter-view complementary information. To address this challenge, this work proposes a multi-scale interaction network that enhances single-view feature representation through a multi-scale spatial-channel attention module and introduces a dual-view epipolar attention module to achieve precise cross-view matching along epipolar lines by incorporating epipolar geometric constraints with an optimal transport algorithm. The method innovatively combines large separable convolutional kernel attention with a geometry-guided optimal transport mechanism. Extensive experiments demonstrate that the proposed approach outperforms most state-of-the-art methods on multiple benchmark datasets, and ablation studies confirm the effectiveness of each component.
This work addresses the entanglement of routing and filtering functions in conventional attention mechanisms, which leads to structural opacity and optimization challenges. The authors propose S-D Attention, which explicitly decouples the interaction matrix into a low-rank routing component and a symmetric filtering component, thereby clearly distinguishing these two mechanisms for the first time. They further uncover that routing self-organizes into a spectral cascade phenomenon in deep networks. Leveraging this insight, they achieve stable training without layer normalization and validate their approach using linear attention variants (e.g., ELU+1) and effective rank analysis. Experiments show that linearizing the first seven layers of a 125M-parameter model incurs less than 5% perplexity degradation, while cascade-informed architectures reduce attention parameters by 47%–65% with only a 3.9%–8.4% increase in perplexity.