Score
Designs, implements, and evaluates neural attention modules and their variants—including softmax vs. linear attention, gated attention, attention maps and masks, and attention biases—focusing on their mathematical behavior, computational complexity, and impact on model accuracy. Builds and integrates attention mechanisms into architectures, models and modifies key–value caching and summarization strategies for bounded memory, and performs comparative analyses and engineering adaptations (e.g., low-resolution maps, parameter-efficient variants, streaming adaptations) to meet accuracy and efficiency constraints.
Existing attention optimization frameworks require extensive manual tuning for heterogeneous hardware and struggle to accommodate novel attention variants or hardware configurations. To address this, we propose a modular, programmable attention computation framework. Our approach introduces: (1) a decomposable attention operator design, enabling flexible composition of arbitrary attention variants; and (2) a unified intermediate representation (IR) coupled with a multi-backend auto-scheduler based on programmable kernel templates, facilitating algorithm-hardware co-optimization. The framework automatically adapts attention implementations across diverse model architectures and hardware platforms without manual re-tuning. Experimental evaluation demonstrates up to 10× speedup over state-of-the-art solutions on non-mainstream hardware configurations, including specialized accelerators and emerging processor architectures. The framework achieves both generality and efficiency while significantly reducing deployment overhead. Open-source implementation is publicly available.
This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
To address excessive computational overhead when deploying vision models on resource-constrained devices, this paper proposes an efficient Vision Transformer (ViT) architecture design framework. Methodologically: (1) it optimizes the input-output data pathway to enhance representational capacity of lightweight models; (2) it restructures the context window of computationally constrained attention mechanisms to improve local-global modeling efficiency; and (3) it leverages the invertibility and explicit probabilistic modeling properties of normalizing flows to enable high-fidelity, low-overhead knowledge distillation. Experiments demonstrate that the proposed approach achieves comparable or superior accuracy on benchmarks such as ImageNet, while requiring significantly fewer parameters and FLOPs. It also substantially reduces inference latency and memory footprint. The framework establishes a scalable new paradigm for efficient visual understanding at the edge.
The quadratic time and memory complexity of Transformer self-attention severely hinders efficient long-context modeling. This paper presents a systematic survey and reconstruction of efficient attention mechanisms for large language models, proposing the first unified taxonomy encompassing both linearization paradigms (e.g., kernel-based approximations and fast weight dynamics) and sparsification paradigms (e.g., fixed patterns, block-wise routing, and clustering-driven selection). It innovatively integrates algorithmic design with hardware-aware optimization, clarifying integration pathways for purely efficient attention and hybrid architectures in large-scale pretraining. Furthermore, it establishes a comprehensive reference framework spanning theoretical analysis, algorithmic implementation, and engineering deployment. The work delivers a systematic design paradigm and practical guidelines for scalable long-context language models.
This work addresses the quadratic computational complexity of standard self-attention in long sequences, which hinders model scalability. The study systematically compares Softmax attention with four recurrent linear attention architectures from the DeltaNet family, unifying their memory mechanisms under a common formulation, and introduces a lightweight Cross-Layer Value Routing (CLVR) mechanism. Experiments at a scale of 350M parameters and 15B tokens reveal that Kimi Delta Attention paired with the Muon optimizer achieves the lowest validation loss, while a pure Gated DeltaNet stack attains the highest training throughput under AdamW. CLVR consistently reduces validation loss across both DeltaNet and its gated variants. This is the first systematic analysis comparing multiple linear attention variants in terms of representational capacity, memory decay characteristics, and training efficiency.
This study investigates differences in neural activation patterns across diverse cognitive tasks among various large language model architectures. Employing a unified framework, the authors systematically analyze final-layer activations, attention entropy, and sparsity across six prominent architectures on twelve task categories, yielding 144 task–model combinations. The work reveals, for the first time, a fundamental distinction between encoder- and decoder-based models in their task-processing mechanisms: mathematical reasoning consistently elicits the highest attention entropy, while decoder-only models exhibit significantly greater activation sparsity. These findings demonstrate the joint influence of architecture type and task category on internal representations, providing empirical guidance for model selection and optimization in large-scale data scenarios.
This work addresses the unclear efficacy of attention mechanisms versus state space models (SSMs) or linear attention components in current hybrid language architectures. To systematically evaluate the necessity and functional division of these modules, the authors propose a functional component ablation framework employing group-wise ablation, layer-wise scanning, positional ablation, and randomized controlled trials across multiple benchmarks. Their analysis provides the first quantitative evidence that SSMs or linear attention serve as the true modeling backbone: their removal degrades perplexity by over 35,000-fold—far exceeding the 82-fold degradation from removing attention. The study further uncovers positional gradient effects, functional redundancy across components, and demonstrates that hybrid models exhibit 20–119 times greater robustness to random layer removal than pure Transformers.
This work addresses the O(n²) memory access bottleneck inherent in the attention mechanism of conventional Transformers. Leveraging the Mathematics of Arrays (MoA) framework, the authors algebraically reformulate scaled dot-product attention and numerically stable softmax, deriving—through purely algebraic means—a Denotational Normal Form (DNF) that eliminates intermediate tensors. This formulation reduces memory access complexity to O(n) and formally guarantees memory optimality via a theorem established prior to code generation. By integrating the Operational Normal Form with dimension-raising hardware mapping techniques, the approach is verified correct in double-precision floating-point arithmetic and projected to achieve 2–100× speedup and 2–50× energy reduction, with performance gains amplifying as problem scale increases.