Score
Designs and implements attention modules that learn or adapt convolutional/filter kernels dynamically to provide selective, multi-scale aggregation across channels and frequencies; builds and analyzes trainable kernel-attention mechanisms including selective-kernel networks, dynamic convolutional kernels, kernel learning and diversification methods, and efficient linear approximations such as RFF-based linear attention.
This work uncovers the fundamental reason why linearized attention mechanisms fail to converge within the Neural Tangent Kernel (NTK) framework and elucidates their dual impact on model performance and robustness. By constructing a linearized attention operator that exactly corresponds to the data-dependent Gram kernel, and integrating NTK theory with spectral analysis, the study establishes—for the first time—a theoretical link between the cubic amplification of the Gram matrix condition number and the required network width (m = Ω(κ⁶)). It introduces the notion of “influence plasticity” to unify the explanation of attention’s expressive power and its vulnerability to adversarial perturbations. Empirical results show that practical training widths fall far below the theoretical threshold, yet exhibit 6–9 times higher influence plasticity than ReLU networks, conferring superior task adaptability at the cost of heightened sensitivity to data poisoning.
This work addresses the lack of theoretical grounding in existing sparse attention mechanisms, whose sparsity patterns and design principles remain largely heuristic. By establishing a formal correspondence between sparse attention and kernel regression with compact support, the study reveals that α-entmax attention (with α = 1 + 1/n) is equivalent to Nadaraya–Watson estimators derived from classical bounded-support kernels such as Epanechnikov and biweight. This equivalence demonstrates that sparsity arises inherently from the bounded support of these kernels, offering a principled theoretical alternative to ad hoc strategies like top-k selection. Leveraging this insight, the authors develop a tunable sparse attention mechanism integrated into the Memory Mosaics architecture, achieving performance on par with state-of-the-art models across language modeling, in-context learning, and length generalization tasks, thereby validating the efficacy and generality of the proposed theoretical framework.
This study addresses the limitations of conventional convolutional neural networks in multi-task scenarios, particularly their weak generalization and poor cross-modal adaptability. Building upon a ResNet-18 backbone, the authors systematically evaluate five dynamic convolution and attention mechanisms—including hard attention, local/global soft attention, and Omni-directional Convolution (ODConv)—across image classification, segmentation, and time series analysis tasks. Experimental results demonstrate that the proposed approaches significantly enhance the model’s adaptive capacity to complex spatial patterns and improve cross-task generalization. Consistent performance gains over standard CNNs are observed on Tiny ImageNet, Pascal VOC, and UCR datasets, with ODConv exhibiting particularly strong performance in complex image-related tasks.
This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)
Existing attention mechanisms operate on discrete sequences, limiting their applicability to continuous function spaces essential for scientific machine learning tasks such as PDE solving and physical simulation. Method: This work generalizes attention to continuous function spaces by introducing the Transformer Neural Operator (TNO), the first rigorously defined attention mechanism on functions. It establishes a mathematically sound formulation of functional attention and proposes a patching-based continuous attention mechanism coupled with an efficient discretization strategy to mitigate computational complexity in high dimensions. Contribution/Results: TNO is proven to be a universal approximator for arbitrary continuous operators. Experiments demonstrate that it significantly outperforms state-of-the-art neural operators across diverse PDE benchmarks and physics-informed simulation tasks, validating its effectiveness, scalability, and generalization capability in scientific machine learning.
This study addresses the integral divergence and model collapse issues arising from exponential attention in Transformers under heavy-tailed distributions. We propose replacing standard softmax with slow-growing kernels and introducing Symlog data preprocessing. By constructing a dedicated benchmark and employing Wasserstein distance metrics, we systematically evaluate the synergistic effects of slow-growing kernels and data transformations on operator learning within post-normalization architectures. Experimental results demonstrate that the proposed method effectively overcomes integral collapse, significantly outperforming conventional softmax attention mechanisms. This work establishes a theoretically grounded and practically valuable new paradigm for operator learning on heavy-tailed data.
研究通过分析高维度下的注意力索引模型,揭示了注意力机制训练动态特性,并提出使用有限截断系统近似无限矩阵矩层次以优化学习过程。
This work addresses a key limitation of existing Transformer-based operator learning methods, which discretize continuous fields into independent tokens and thereby neglect the global structure of function spaces, hindering their ability to model mappings between infinite-dimensional functions. To overcome this, the authors propose Functional Attention—a novel mechanism inspired by geometric functional maps—that generalizes attention from pointwise affinities to linear functional correspondences over function spaces. By replacing the conventional softmax with a structured linear operator and integrating adaptive basis construction, the method explicitly captures global dependencies. The resulting representation is compact, generalizable, and resolution-invariant, achieving state-of-the-art performance on tasks such as partial differential equation solving, 3D segmentation, and regression, while demonstrating strong robustness across diverse discretization schemes.
This work addresses the limitations of conventional linear attention mechanisms, which suffer from approximation-induced errors, gradient explosion, and attention dilution. To overcome these issues, the authors propose a linear-complexity attention mechanism that eliminates approximation error by introducing novel kernel functions—such as the Hadamard Exp kernel and the squared Euclidean distance kernel—that satisfy non-negativity, discriminability, and geometric interpretability, enabling exact kernel decomposition. Furthermore, the model incorporates a Hyper Link structure, a Memory Lobe module, and a Mixture-of-Experts routing bias mechanism to enhance memory capacity, semantic alignment, and training stability. The resulting approach achieves efficient and accurate linear attention computation while preserving model performance and effectively mitigating gradient degradation and attention dilution.
To address the trade-off between the quadratic complexity of soft attention and the accuracy degradation of existing linear attention methods—caused by fixed, non-adaptive random feature mappings—this paper proposes LUNA, a linear attention mechanism with learnable kernel feature mappings. Its core innovation lies in parameterizing the feature mapping of a kernel function and optimizing it end-to-end, enabling the first adaptive learning of feature mappings within kernelized linear attention. LUNA preserves O(n) time and space complexity while substantially enhancing modeling capacity. Theoretically, it guarantees positive-definiteness of the induced kernel and ensures generalization bounds, supporting efficient streaming inference. Experiments demonstrate that LUNA achieves state-of-the-art average accuracy on Long Range Arena and significantly outperforms fixed-mapping baselines when applied as a drop-in replacement in post-hoc BERT and ViT attention layers—nearly fully recovering the performance of the original models.