Score
Design and implement neural network modules that adaptively reweight or recalibrate feature channels — including per-instance channel attention, class-aware scaling, and gating mechanisms — so that discriminative channels are amplified and irrelevant channels are suppressed. Build and analyze the resulting channel-wise or class-specific attention maps and evaluate their effects on feature discriminability, localization of fine-grained regions, and performance for rare/tail classes.
In few-shot learning, deep networks suffer from channel bias—over-reliance on source-task-discriminative channels—leading to feature redundancy: most channels exhibit high intra-class variance and low inter-class separability, thereby hindering adaptation to novel tasks. Method: We establish, for the first time, the causal chain “channel bias → feature redundancy → performance degradation,” theoretically proving that redundancy is exacerbated under low-data regimes. We reveal the “less-is-more” principle: retaining only 1–5% of the most discriminative channels significantly improves accuracy, while >95% contribute negatively—a detrimental effect that attenuates with increasing sample size. Accordingly, we propose Augmented Feature Importance Adjustment (AFIA), a theoretically grounded method integrating data augmentation and soft channel masking to suppress redundancy. Results: AFIA achieves consistent and significant improvements across standard few-shot benchmarks, offering both a new theoretical principle and a practical tool for few-shot learning.
High redundancy in deep networks impedes the simultaneous achievement of efficiency and accuracy. To address this, we propose the Partial Channel Mechanism (PCM), which dynamically partitions feature map channels and assigns heterogeneous operations—such as convolution, attention, and pooling—to distinct channel subsets, enabling fine-grained computational resource allocation. Based on PCM, we design Partial Attention Convolution (PATConv) and Dynamic Partial Convolution (DPConv), and introduce PartialNet—the first hybrid network family leveraging channel-wise sparsity. PartialNet supports learnable, adaptive channel partitioning ratios, substantially reducing both parameter count and FLOPs. On ImageNet-1K, it surpasses multiple state-of-the-art models in both top-1 accuracy and inference speed. Moreover, it achieves competitive performance on COCO object detection and instance segmentation tasks, demonstrating broad applicability and effectiveness across vision domains.
This work addresses the lack of systematic analysis and unified design principles in existing channel-spatial attention fusion strategies. Under a consistent experimental framework, the authors construct and comprehensively evaluate 18 channel-spatial attention topologies, spanning serial, parallel, multi-scale, and residual architectures. Extensive experiments across diverse vision and medical imaging datasets reveal a coupling relationship among data scale, architectural design, and performance. The study proposes practical guidelines for attention module construction tailored to data regime size: cascaded channel–multi-scale spatial attention excels in small-sample tasks; learnable parallel fusion achieves optimal results at medium scales; and large-scale scenarios benefit from parallel structures augmented with dynamic gating. Additionally, the work validates the advantage of spatial-before-channel ordering for fine-grained classification and demonstrates the efficacy of residual connections in mitigating gradient vanishing.
Existing semi-dense matching methods apply uniform weighting to all pixels during attention-based feature extraction, making them susceptible to noise from redundant and irrelevant regions and yielding poorly discriminative attention weights. To address this, we propose a matchability-aware dual-path dynamic reweighting mechanism: (i) injecting learnable matchability bias into attention logits, and (ii) performing matchability-driven post-attention adaptive scaling of value features. This mechanism employs a lightweight binary classification head to estimate pixel-wise matchability in real time, enabling fine-grained, semantics-aware attention modulation. Fully embedded within the Transformer architecture, our method requires no additional supervision or pretraining. Evaluated on HPatches, ETH3D, and SUN3D—three major benchmarks—it consistently outperforms state-of-the-art methods, achieving significant improvements in repeatability, matching accuracy, and robustness to viewpoint and illumination variations.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
为解决细粒度视觉识别中通道注意力机制效果有限的问题,提出ConCA方法,通过结合均值和负输入熵形成双重描述符,并使用深度可分离1-D卷积多层感知器映射到每通道权重。
Convolutional neural networks often exhibit poor generalization and fairness issues due to their reliance on spurious correlations in training data. This work identifies global average pooling as a key factor that entangles core features with spurious ones during feature aggregation. To address this, the authors propose a retrainable attention-based aggregation module as a post-processing step, which adaptively weights spatial locations prior to aggregation to selectively suppress spurious features. The method jointly optimizes the classification head and feature aggregation without requiring modifications to the backbone network. Experimental results demonstrate that the approach significantly outperforms existing Debiased Feature Reweighting (DFR) methods across multiple datasets and evaluation metrics, effectively reducing the model’s dependence on spurious correlations.
This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.
This work addresses a critical limitation in existing channel pruning methods, which conflate two orthogonal dimensions—task relevance and local substitutability—thereby constraining performance. For the first time, this study explicitly disentangles these concepts: task relevance quantifies a channel’s contribution to the target objective, while local substitutability measures whether its function can be compensated by other channels within the same layer. Theoretical analysis and empirical evidence demonstrate that these two properties rapidly decouple during training, with local substitutability emerging as a more reliable criterion for pruning. Through comprehensive validation—including input attribution, channel overlap analysis, task information metrics, residual gradient examination, and ablation studies—this approach consistently outperforms conventional pruning strategies across multiple architectures and benchmarks, including CIFAR-100 and ImageNet.
This work elucidates the underlying mechanism of gated MLPs by offering the first explanation of their success through the lens of symmetry breaking. It demonstrates that a gated MLP can be interpreted as a rank-1 approximation of bilinear attention, where the query and key correspond to two distinct factors, and the nonlinear activation is applied exclusively to one factor. This asymmetric treatment breaks both the exchange symmetry between the two factors and the inverse scaling symmetry induced by non-homogeneous activation functions. The analysis establishes a theoretical connection between gated MLPs and attention mechanisms, clarifying the origin of their performance advantages and providing a principled foundation for designing novel, efficient architectures.